{
  "id": 209622,
  "title": "Public/Private 13th solution (team takoi + kurupical)",
  "url": "/competitions/riiid-test-answer-prediction/writeups/takoi-kurupical-public-private-13th-solution-team-",
  "author_name": "",
  "post_date": "2021-01-09T10:26:59.530Z",
  "votes": 65,
  "comment_count": 14,
  "views": 0,
  "content": "<p>First of all, thanks to Hoon Pyo (Tim) Jeon and Kaggle team for such an interesting competition.<br>\nAnd congratulates to all the winning teams!</p>\n<p>The following is the team takoi + kurupical solution.</p>\n<p><br><br>\n<br></p>\n<h1>Team takoi + kurupical Overview</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1146523%2F3f2943b27feb7d884472c201cd032d5b%2F1.jpg?generation=1610075778291231&amp;alt=media\" alt=\"\"></p>\n<h1>validation</h1>\n<p><a href=\"https://www.kaggle.com/tito\" target=\"_blank\">@tito</a>'s validation strategy. <a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">https://www.kaggle.com/its7171/cv-strategy</a></p>\n<p><br><br>\n<br></p>\n<h1><strong>kurupical side</strong></h1>\n<p>I would like to have three kaggler to thank.<br>\n<a href=\"https://www.kaggle.com/takoi\" target=\"_blank\">@takoi</a> for inviting me to form a team. If it weren't for you, I couldn't reach this rank!<br>\n<a href=\"https://www.kaggle.com/limerobot\" target=\"_blank\">@limerobot</a> for sharing DSB 3rd solution. I'm beginner in transformer for time-series data, so I learned a lot from your solution!<br>\n<a href=\"https://www.kaggle.com/wangsg\" target=\"_blank\">@wangsg</a> for sharing notebook <a href=\"https://www.kaggle.com/wangsg/a-self-attentive-model-for-knowledge-tracing.!\" target=\"_blank\">https://www.kaggle.com/wangsg/a-self-attentive-model-for-knowledge-tracing.!</a><br>\nI used this notebook as a baseline and finally get 0.809 CV for single transformer.</p>\n<h2>model</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1146523%2F411c8ee779411bb443284b572980ef21%2F2.jpg?generation=1610075796425059&amp;alt=media\" alt=\"\"></p>\n<h2>hyper parameters</h2>\n<ul>\n<li>20epochs</li>\n<li>AdamW(lr=1e-3, weight_decay=0.1)</li>\n<li>linear_with_warmup(lr=1e-3, warmup_epoch=2)</li>\n</ul>\n<h2>worked for me</h2>\n<ul>\n<li>baseline (SAKT, <a href=\"https://www.kaggle.com/wangsg/a-self-attentive-model-for-knowledge-tracing\" target=\"_blank\">https://www.kaggle.com/wangsg/a-self-attentive-model-for-knowledge-tracing</a>)</li>\n<li>use all data (this notebook use only last 100 history per user)</li>\n<li>embedding concat (not add) and Linear layer after cat embedding(@limerobot DSB2019 3rd solution) (+0.03) </li>\n<li>Add min(timestamp_delta//1000, 300) (+0.02)</li>\n<li>Add \"index that user answered same content_id at last\" (+0.005)</li>\n<li>Transformer Encoder n_layers 2 -&gt; 4 (+0.002)</li>\n<li>weight_decay 0.01 -&gt; 0.1 (+0.002)</li>\n<li>LIT structure in EncoderLayer (+0.002)</li>\n</ul>\n<h2>not worked for me</h2>\n<p>I did over 300 experiments, and only about 20 of them were successful.</p>\n<ul>\n<li>SAINT structure (Transformer Encoder/Decoder)</li>\n<li>Positional Encoding</li>\n<li>Consider timeseries<ul>\n<li>timedelta.cumsum() / timedelta.sum()</li>\n<li>np.log10(timedelta.cumsum()).astype(int) as category feature and embedding<br>\netc…</li></ul></li>\n<li>optimizer AdaBelief, LookAhead(Adam), RAdam</li>\n<li>more n_layers(4 =&gt; 6), more embedding_dimention (256 =&gt; 512)</li>\n<li>output only the end of the sequence</li>\n<li>large binning for elapsed_time/timedelta (500, 1000, etc…)</li>\n<li>treat elapsed_time and timedelta as continuous </li>\n</ul>\n<p><br><br>\n<br></p>\n<h1><strong>takoi side</strong></h1>\n<p>I made 1 LightGBM and 8 NN models. The model that combined Transformer and LSTM had the best CV. Here is architecture and brief description.<br>\n​</p>\n<h2>Transformer + LSTM</h2>\n<p>​<br>\n<br><br>\n​<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1146523%2Fcab7462101930cf01ccbdd867e3b8a68%2F2%20(1).jpeg?generation=1610097396888437&amp;alt=media\" alt=\"\"></p>\n<h3>features</h3>\n<p>I used 17 features. 15 features were computed per user_id. 2 features were computed per content_id.</p>\n<h4>main features</h4>\n<ul>\n<li>sum of answered correctly</li>\n<li>average of answered correctly</li>\n<li>sum of answered correctly for tag_user_id</li>\n<li>average of answered correctly for tag_user_id</li>\n<li>lag time</li>\n<li>lag time of same content_id</li>\n<li>previous answered correctly for the same content_id</li>\n<li>distance between the same content_id</li>\n<li>average of answered correctly for each content_id</li>\n<li>average of lag time for each content_id<br>\n​<br>\n<br><br>\n​</li>\n</ul>\n<h2>LightGBM</h2>\n<p>I used 97 features. The following are the main features.</p>\n<ul>\n<li>sum of answered correctly</li>\n<li>average of answered correctly</li>\n<li>sum of answered correctly for tag_user_id</li>\n<li>average of answered correctly for tag_user_id</li>\n<li>lag time</li>\n<li>lag time of same part</li>\n<li>lag time of same content_id</li>\n<li>previous answered correctly for the same content_id</li>\n<li>distance between the same content_id</li>\n<li>Word2Vec features of content_id</li>\n<li>decayed features (average of answered correctly)</li>\n<li>number of consecutive times with the same user answer</li>\n<li>flag if the user answer being answered in succession matches the correct answer</li>\n<li>average of answered correctly for each content_id</li>\n<li>average of lag time for each content_id</li>\n</ul>",
  "messages": [
    {
      "id": "1143748",
      "postDate": "01/08/2021 03:30:33",
      "content": "<p>First of all, thanks to Hoon Pyo (Tim) Jeon and Kaggle team for such an interesting competition.<br>\nAnd congratulates to all the winning teams!</p>\n<p>The following is the team takoi + kurupical solution.</p>\n<p><br><br>\n<br></p>\n<h1>Team takoi + kurupical Overview</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1146523%2F3f2943b27feb7d884472c201cd032d5b%2F1.jpg?generation=1610075778291231&amp;alt=media\" alt=\"\"></p>\n<h1>validation</h1>\n<p><a href=\"https://www.kaggle.com/tito\" target=\"_blank\">@tito</a>'s validation strategy. <a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">https://www.kaggle.com/its7171/cv-strategy</a></p>\n<p><br><br>\n<br></p>\n<h1><strong>kurupical side</strong></h1>\n<p>I would like to have three kaggler to thank.<br>\n<a href=\"https://www.kaggle.com/takoi\" target=\"_blank\">@takoi</a> for inviting me to form a team. If it weren't for you, I couldn't reach this rank!<br>\n<a href=\"https://www.kaggle.com/limerobot\" target=\"_blank\">@limerobot</a> for sharing DSB 3rd solution. I'm beginner in transformer for time-series data, so I learned a lot from your solution!<br>\n<a href=\"https://www.kaggle.com/wangsg\" target=\"_blank\">@wangsg</a> for sharing notebook <a href=\"https://www.kaggle.com/wangsg/a-self-attentive-model-for-knowledge-tracing.!\" target=\"_blank\">https://www.kaggle.com/wangsg/a-self-attentive-model-for-knowledge-tracing.!</a><br>\nI used this notebook as a baseline and finally get 0.809 CV for single transformer.</p>\n<h2>model</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1146523%2F411c8ee779411bb443284b572980ef21%2F2.jpg?generation=1610075796425059&amp;alt=media\" alt=\"\"></p>\n<h2>hyper parameters</h2>\n<ul>\n<li>20epochs</li>\n<li>AdamW(lr=1e-3, weight_decay=0.1)</li>\n<li>linear_with_warmup(lr=1e-3, warmup_epoch=2)</li>\n</ul>\n<h2>worked for me</h2>\n<ul>\n<li>baseline (SAKT, <a href=\"https://www.kaggle.com/wangsg/a-self-attentive-model-for-knowledge-tracing\" target=\"_blank\">https://www.kaggle.com/wangsg/a-self-attentive-model-for-knowledge-tracing</a>)</li>\n<li>use all data (this notebook use only last 100 history per user)</li>\n<li>embedding concat (not add) and Linear layer after cat embedding(@limerobot DSB2019 3rd solution) (+0.03) </li>\n<li>Add min(timestamp_delta//1000, 300) (+0.02)</li>\n<li>Add \"index that user answered same content_id at last\" (+0.005)</li>\n<li>Transformer Encoder n_layers 2 -&gt; 4 (+0.002)</li>\n<li>weight_decay 0.01 -&gt; 0.1 (+0.002)</li>\n<li>LIT structure in EncoderLayer (+0.002)</li>\n</ul>\n<h2>not worked for me</h2>\n<p>I did over 300 experiments, and only about 20 of them were successful.</p>\n<ul>\n<li>SAINT structure (Transformer Encoder/Decoder)</li>\n<li>Positional Encoding</li>\n<li>Consider timeseries<ul>\n<li>timedelta.cumsum() / timedelta.sum()</li>\n<li>np.log10(timedelta.cumsum()).astype(int) as category feature and embedding<br>\netc…</li></ul></li>\n<li>optimizer AdaBelief, LookAhead(Adam), RAdam</li>\n<li>more n_layers(4 =&gt; 6), more embedding_dimention (256 =&gt; 512)</li>\n<li>output only the end of the sequence</li>\n<li>large binning for elapsed_time/timedelta (500, 1000, etc…)</li>\n<li>treat elapsed_time and timedelta as continuous </li>\n</ul>\n<p><br><br>\n<br></p>\n<h1><strong>takoi side</strong></h1>\n<p>I made 1 LightGBM and 8 NN models. The model that combined Transformer and LSTM had the best CV. Here is architecture and brief description.<br>\n​</p>\n<h2>Transformer + LSTM</h2>\n<p>​<br>\n<br><br>\n​<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1146523%2Fcab7462101930cf01ccbdd867e3b8a68%2F2%20(1).jpeg?generation=1610097396888437&amp;alt=media\" alt=\"\"></p>\n<h3>features</h3>\n<p>I used 17 features. 15 features were computed per user_id. 2 features were computed per content_id.</p>\n<h4>main features</h4>\n<ul>\n<li>sum of answered correctly</li>\n<li>average of answered correctly</li>\n<li>sum of answered correctly for tag_user_id</li>\n<li>average of answered correctly for tag_user_id</li>\n<li>lag time</li>\n<li>lag time of same content_id</li>\n<li>previous answered correctly for the same content_id</li>\n<li>distance between the same content_id</li>\n<li>average of answered correctly for each content_id</li>\n<li>average of lag time for each content_id<br>\n​<br>\n<br><br>\n​</li>\n</ul>\n<h2>LightGBM</h2>\n<p>I used 97 features. The following are the main features.</p>\n<ul>\n<li>sum of answered correctly</li>\n<li>average of answered correctly</li>\n<li>sum of answered correctly for tag_user_id</li>\n<li>average of answered correctly for tag_user_id</li>\n<li>lag time</li>\n<li>lag time of same part</li>\n<li>lag time of same content_id</li>\n<li>previous answered correctly for the same content_id</li>\n<li>distance between the same content_id</li>\n<li>Word2Vec features of content_id</li>\n<li>decayed features (average of answered correctly)</li>\n<li>number of consecutive times with the same user answer</li>\n<li>flag if the user answer being answered in succession matches the correct answer</li>\n<li>average of answered correctly for each content_id</li>\n<li>average of lag time for each content_id</li>\n</ul>",
      "rawMarkdown": "First of all, thanks to Hoon Pyo (Tim) Jeon and Kaggle team for such an interesting competition.\nAnd congratulates to all the winning teams!\n\nThe following is the team takoi + kurupical solution.\n\n<br>\n<br>\n\n# Team takoi + kurupical Overview\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1146523%2F3f2943b27feb7d884472c201cd032d5b%2F1.jpg?generation=1610075778291231&alt=media)\n\n# validation\n@tito's validation strategy. https://www.kaggle.com/its7171/cv-strategy\n\n<br>\n<br>\n# **kurupical side**\nI would like to have three kaggler to thank.\n@takoi for inviting me to form a team. If it weren't for you, I couldn't reach this rank!\n@limerobot for sharing DSB 3rd solution. I'm beginner in transformer for time-series data, so I learned a lot from your solution!\n@wangsg for sharing notebook https://www.kaggle.com/wangsg/a-self-attentive-model-for-knowledge-tracing.!\nI used this notebook as a baseline and finally get 0.809 CV for single transformer.\n\n## model\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1146523%2F411c8ee779411bb443284b572980ef21%2F2.jpg?generation=1610075796425059&alt=media)\n\n## hyper parameters\n* 20epochs\n* AdamW(lr=1e-3, weight_decay=0.1)\n* linear_with_warmup(lr=1e-3, warmup_epoch=2)\n\n## worked for me\n* baseline (SAKT, https://www.kaggle.com/wangsg/a-self-attentive-model-for-knowledge-tracing)\n* use all data (this notebook use only last 100 history per user)\n* embedding concat (not add) and Linear layer after cat embedding(@limerobot DSB2019 3rd solution) (+0.03) \n* Add min(timestamp_delta//1000, 300) (+0.02)\n* Add \"index that user answered same content_id at last\" (+0.005)\n* Transformer Encoder n_layers 2 -> 4 (+0.002)\n* weight_decay 0.01 -> 0.1 (+0.002)\n* LIT structure in EncoderLayer (+0.002)\n\n## not worked for me\nI did over 300 experiments, and only about 20 of them were successful.\n\n* SAINT structure (Transformer Encoder/Decoder)\n* Positional Encoding\n* Consider timeseries\n  * timedelta.cumsum() / timedelta.sum()\n  * np.log10(timedelta.cumsum()).astype(int) as category feature and embedding\n  etc...\n* optimizer AdaBelief, LookAhead(Adam), RAdam\n* more n_layers(4 => 6), more embedding_dimention (256 => 512)\n* output only the end of the sequence\n* large binning for elapsed_time/timedelta (500, 1000, etc...)\n* treat elapsed_time and timedelta as continuous \n\n<br>\n<br>\n\n# **takoi side**\nI made 1 LightGBM and 8 NN models. The model that combined Transformer and LSTM had the best CV. Here is architecture and brief description.\n​\n## Transformer + LSTM\n​\n</br>\n​\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1146523%2Fcab7462101930cf01ccbdd867e3b8a68%2F2%20(1).jpeg?generation=1610097396888437&alt=media)\n\n### features\nI used 17 features. 15 features were computed per user_id. 2 features were computed per content_id.\n#### main features\n - sum of answered correctly\n - average of answered correctly\n - sum of answered correctly for tag_user_id\n - average of answered correctly for tag_user_id\n - lag time\n - lag time of same content_id\n - previous answered correctly for the same content_id\n - distance between the same content_id\n - average of answered correctly for each content_id\n - average of lag time for each content_id\n​\n</br>\n​\n## LightGBM\nI used 97 features. The following are the main features.\n- sum of answered correctly\n- average of answered correctly\n- sum of answered correctly for tag_user_id\n- average of answered correctly for tag_user_id\n- lag time\n- lag time of same part\n- lag time of same content_id\n- previous answered correctly for the same content_id\n- distance between the same content_id\n- Word2Vec features of content_id\n- decayed features (average of answered correctly)\n- number of consecutive times with the same user answer\n- flag if the user answer being answered in succession matches the correct answer\n- average of answered correctly for each content_id\n- average of lag time for each content_id",
      "votes": null
    },
    {
      "id": "1143785",
      "postDate": "01/08/2021 04:15:57",
      "content": "<p>Congrats on 13th place <a href=\"https://www.kaggle.com/kurupical\" target=\"_blank\">@kurupical</a> and <a href=\"https://www.kaggle.com/takoihiraokazu\" target=\"_blank\">@takoihiraokazu</a> and thanks for sharing your solution</p>",
      "rawMarkdown": "Congrats on 13th place @kurupical and @takoihiraokazu and thanks for sharing your solution",
      "votes": null
    },
    {
      "id": "1144389",
      "postDate": "01/08/2021 12:39:23",
      "content": "<p>Very nice write-up, thanks for sharing ! And congratz on the gold finish!</p>",
      "rawMarkdown": "Very nice write-up, thanks for sharing ! And congratz on the gold finish!",
      "votes": null
    },
    {
      "id": "1145474",
      "postDate": "01/09/2021 07:04:38",
      "content": "<p>congratulations！</p>",
      "rawMarkdown": "congratulations！",
      "votes": null
    },
    {
      "id": "1145485",
      "postDate": "01/09/2021 07:13:36",
      "content": "<p>Congratulations on a cool finish!</p>\n<ul>\n<li>What's same_question_index and target?</li>\n<li>distance between the same content_id -&gt; so this is the difference b/w the sequence no when a particular content_id was last seen Vs their current sequence number assigned, right?</li>\n<li>decayed features (average of answered correctly) -&gt; How did you decay them? Did you do something like <code>alpha*new_value + (1-alpha)*old_value</code>?</li>\n<li>LIT structure in EncoderLayer. What's LIT?</li>\n</ul>\n<p>Thanks a lot!</p>",
      "rawMarkdown": "Congratulations on a cool finish!\n\n- What's same_question_index and target?\n- distance between the same content_id -> so this is the difference b/w the sequence no when a particular content_id was last seen Vs their current sequence number assigned, right?\n- decayed features (average of answered correctly) -> How did you decay them? Did you do something like `alpha*new_value + (1-alpha)*old_value`?\n- LIT structure in EncoderLayer. What's LIT?\n\nThanks a lot!",
      "votes": null
    },
    {
      "id": "1145614",
      "postDate": "01/09/2021 08:57:14",
      "content": "<p>Thank you!<br>\nI answer question about my side.</p>\n<pre><code>distance between the same content_id -&gt; so this is the difference b/w the sequence no when a particular content_id was last seen Vs their current sequence number assigned, right?\n</code></pre>\n<p>Yes.</p>\n<pre><code>decayed features (average of answered correctly) -&gt; How did you decay them? Did you do something like alpha*new_value + (1-alpha)*old_value?\n</code></pre>\n<p>I decayed the value as <code>0.3 * new_vlaue + (1 - 0.3) * old_value</code> .</p>",
      "rawMarkdown": "Thank you!\nI answer question about my side.\n\n``` \ndistance between the same content_id -> so this is the difference b/w the sequence no when a particular content_id was last seen Vs their current sequence number assigned, right?\n```\nYes.\n\n```\ndecayed features (average of answered correctly) -> How did you decay them? Did you do something like alpha*new_value + (1-alpha)*old_value?\n```\nI decayed the value as ``` 0.3 * new_vlaue + (1 - 0.3) * old_value``` .",
      "votes": null
    },
    {
      "id": "1145626",
      "postDate": "01/09/2021 09:07:23",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> !</p>\n<blockquote>\n  <p>What's same_question_index and target?</p>\n</blockquote>\n<p>-&gt; same as takoi's distance between the same content_id</p>\n<blockquote>\n  <p>LIT structure in EncoderLayer. What's LIT?</p>\n</blockquote>\n<p>-&gt; Sorry, I forget to explain detail. <a href=\"https://arxiv.org/pdf/2012.14164.pdf\" target=\"_blank\">https://arxiv.org/pdf/2012.14164.pdf</a>. My implements is below. TransformerEncoder is copy from torch.nn.TransformerEncoder, I just replaced the linear part of src2 with a LITLayer.</p>\n<pre><code>class LITLayer(nn.Module):\n    \"\"\"\n    https://arxiv.org/pdf/2012.14164.pdf\n    \"\"\"\n    def __init__(self, input_dim, embed_dim, dropout, activation):\n        super(LITLayer, self).__init__()\n        self.input_dim = input_dim\n        self.embed_dim = embed_dim\n        self.activation = activation\n\n        self.linear1 = nn.Linear(input_dim, embed_dim)\n        self.lstm = nn.LSTM(embed_dim, embed_dim)\n        self.linear2 = nn.Linear(embed_dim, input_dim)\n\n        self.norm_lstm = nn.LayerNorm(embed_dim)\n\n        self.dropout = nn.Dropout(dropout)\n\n    def forward(self, x):\n        x = self.linear1(x)\n        x = self.activation(x)\n        x = self.dropout(x)\n        x, _ = self.lstm(x)\n        x = self.norm_lstm(x)\n        x = self.dropout(x)\n        x = self.linear2(x)\n\n        return x\n\nclass TransformerEncoderLayer(nn.Module):\n    r\"\"\"TransformerEncoderLayer is made up of self-attn and feedforward network.\n    This standard encoder layer is based on the paper \"Attention Is All You Need\".\n    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez,\n    Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in\n    Neural Information Processing Systems, pages 6000-6010. Users may modify or implement\n    in a different way during application.\n\n    Args:\n        d_model: the number of expected features in the input (required).\n        nhead: the number of heads in the multiheadattention models (required).\n        dim_feedforward: the dimension of the feedforward network model (default=2048).\n        dropout: the dropout value (default=0.1).\n        activation: the activation function of intermediate layer, relu or gelu (default=relu).\n\n    Examples::\n        &gt;&gt;&gt; encoder_layer = nn.TransformerEncoderLayer(d_model=512, nhead=8)\n        &gt;&gt;&gt; src = torch.rand(10, 32, 512)\n        &gt;&gt;&gt; out = encoder_layer(src)\n    \"\"\"\n\n    def __init__(self, d_model, nhead, dim_feedforward=256, dropout=0.1, activation=\"relu\"):\n        super(TransformerEncoderLayer, self).__init__()\n        self.self_attn = nn.MultiheadAttention(d_model, nhead, dropout=dropout)\n        # Implementation of Feedforward model\n\n        self.norm1 = nn.LayerNorm(d_model)\n        self.norm2 = nn.LayerNorm(d_model)\n        self.dropout1 = nn.Dropout(dropout)\n        self.dropout2 = nn.Dropout(dropout)\n\n        self.activation = _get_activation_fn(activation)\n\n        self.lit_layer = LITLayer(input_dim=d_model, embed_dim=dim_feedforward, dropout=dropout, activation=self.activation)\n\n    def __setstate__(self, state):\n        if 'activation' not in state:\n            state['activation'] = nn.ReLU\n        super(TransformerEncoderLayer, self).__setstate__(state)\n\n    def forward(self, src: torch.Tensor, src_mask: Optional[torch.Tensor] = None, src_key_padding_mask: Optional[torch.Tensor] = None) -&gt; torch.Tensor:\n        r\"\"\"Pass the input through the encoder layer.\n\n        Args:\n            src: the sequence to the encoder layer (required).\n            src_mask: the mask for the src sequence (optional).\n            src_key_padding_mask: the mask for the src keys per batch (optional).\n\n        Shape:\n            see the docs in Transformer class.\n        \"\"\"\n        src2 = self.self_attn(src, src, src, attn_mask=src_mask,\n                              key_padding_mask=src_key_padding_mask)[0]\n        src = src + self.dropout1(src2)\n        src = self.norm1(src)\n       # src2 = self.linear2(self.dropout(self.activation(self.linear1(src))))\n        src2 = self.lit_layer(src)\n        src = src + self.dropout2(src2)\n        src = self.norm2(src)\n        return src\n</code></pre>",
      "rawMarkdown": "Hi @adityaecdrid !\n\n> What's same_question_index and target?\n\n-> same as takoi's distance between the same content_id\n\n> LIT structure in EncoderLayer. What's LIT?\n\n-> Sorry, I forget to explain detail. https://arxiv.org/pdf/2012.14164.pdf. My implements is below. TransformerEncoder is copy from torch.nn.TransformerEncoder, I just replaced the linear part of src2 with a LITLayer.\n\n```\nclass LITLayer(nn.Module):\n    \"\"\"\n    https://arxiv.org/pdf/2012.14164.pdf\n    \"\"\"\n    def __init__(self, input_dim, embed_dim, dropout, activation):\n        super(LITLayer, self).__init__()\n        self.input_dim = input_dim\n        self.embed_dim = embed_dim\n        self.activation = activation\n\n        self.linear1 = nn.Linear(input_dim, embed_dim)\n        self.lstm = nn.LSTM(embed_dim, embed_dim)\n        self.linear2 = nn.Linear(embed_dim, input_dim)\n\n        self.norm_lstm = nn.LayerNorm(embed_dim)\n\n        self.dropout = nn.Dropout(dropout)\n\n    def forward(self, x):\n        x = self.linear1(x)\n        x = self.activation(x)\n        x = self.dropout(x)\n        x, _ = self.lstm(x)\n        x = self.norm_lstm(x)\n        x = self.dropout(x)\n        x = self.linear2(x)\n\n        return x\n\nclass TransformerEncoderLayer(nn.Module):\n    r\"\"\"TransformerEncoderLayer is made up of self-attn and feedforward network.\n    This standard encoder layer is based on the paper \"Attention Is All You Need\".\n    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez,\n    Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in\n    Neural Information Processing Systems, pages 6000-6010. Users may modify or implement\n    in a different way during application.\n\n    Args:\n        d_model: the number of expected features in the input (required).\n        nhead: the number of heads in the multiheadattention models (required).\n        dim_feedforward: the dimension of the feedforward network model (default=2048).\n        dropout: the dropout value (default=0.1).\n        activation: the activation function of intermediate layer, relu or gelu (default=relu).\n\n    Examples::\n        >>> encoder_layer = nn.TransformerEncoderLayer(d_model=512, nhead=8)\n        >>> src = torch.rand(10, 32, 512)\n        >>> out = encoder_layer(src)\n    \"\"\"\n\n    def __init__(self, d_model, nhead, dim_feedforward=256, dropout=0.1, activation=\"relu\"):\n        super(TransformerEncoderLayer, self).__init__()\n        self.self_attn = nn.MultiheadAttention(d_model, nhead, dropout=dropout)\n        # Implementation of Feedforward model\n\n        self.norm1 = nn.LayerNorm(d_model)\n        self.norm2 = nn.LayerNorm(d_model)\n        self.dropout1 = nn.Dropout(dropout)\n        self.dropout2 = nn.Dropout(dropout)\n\n        self.activation = _get_activation_fn(activation)\n\n        self.lit_layer = LITLayer(input_dim=d_model, embed_dim=dim_feedforward, dropout=dropout, activation=self.activation)\n\n    def __setstate__(self, state):\n        if 'activation' not in state:\n            state['activation'] = nn.ReLU\n        super(TransformerEncoderLayer, self).__setstate__(state)\n\n    def forward(self, src: torch.Tensor, src_mask: Optional[torch.Tensor] = None, src_key_padding_mask: Optional[torch.Tensor] = None) -> torch.Tensor:\n        r\"\"\"Pass the input through the encoder layer.\n\n        Args:\n            src: the sequence to the encoder layer (required).\n            src_mask: the mask for the src sequence (optional).\n            src_key_padding_mask: the mask for the src keys per batch (optional).\n\n        Shape:\n            see the docs in Transformer class.\n        \"\"\"\n        src2 = self.self_attn(src, src, src, attn_mask=src_mask,\n                              key_padding_mask=src_key_padding_mask)[0]\n        src = src + self.dropout1(src2)\n        src = self.norm1(src)\n       # src2 = self.linear2(self.dropout(self.activation(self.linear1(src))))\n        src2 = self.lit_layer(src)\n        src = src + self.dropout2(src2)\n        src = self.norm2(src)\n        return src\n```",
      "votes": null
    },
    {
      "id": "1145637",
      "postDate": "01/09/2021 09:14:36",
      "content": "<p>Thanks guys for the quick response!</p>",
      "rawMarkdown": "Thanks guys for the quick response!",
      "votes": null
    },
    {
      "id": "1146884",
      "postDate": "01/10/2021 06:23:05",
      "content": "<p>Thank you for sharing a great solution and congratulations on gold medal!</p>\n<p>It might be a dumb question but what makes you stack a LSTM layer on a Transformer layer?<br>\nI am studying NN and it would be a great help if you guide me to understand how LSTM is superior to Transformer.</p>\n<p>And, how did you decide which features put in transformer or dense layer? </p>",
      "rawMarkdown": "Thank you for sharing a great solution and congratulations on gold medal!\n\nIt might be a dumb question but what makes you stack a LSTM layer on a Transformer layer?\nI am studying NN and it would be a great help if you guide me to understand how LSTM is superior to Transformer.\n\nAnd, how did you decide which features put in transformer or dense layer?",
      "votes": null
    },
    {
      "id": "1146914",
      "postDate": "01/10/2021 06:52:20",
      "content": "<p>Thanks!</p>\n<pre><code>It might be a dumb question but what makes you stack a LSTM layer on a Transformer layer?\nI am studying NN and it would be a great help if you guide me to understand how LSTM is superior to Transformer.\n</code></pre>\n<p>I think LSTM is better at capturing nearby information than Transformer, but Transformer is better at capturing distant information. That's why I included both.</p>\n<pre><code>And, how did you decide which features put in transformer or dense layer?\n</code></pre>\n<p>I included cumulative and average features for all previous problems of the user, and features that greatly increased the score of LightGBM.</p>",
      "rawMarkdown": "Thanks!\n```\nIt might be a dumb question but what makes you stack a LSTM layer on a Transformer layer?\nI am studying NN and it would be a great help if you guide me to understand how LSTM is superior to Transformer.\n```\nI think LSTM is better at capturing nearby information than Transformer, but Transformer is better at capturing distant information. That's why I included both.\n\n```\nAnd, how did you decide which features put in transformer or dense layer?\n```\nI included cumulative and average features for all previous problems of the user, and features that greatly increased the score of LightGBM.",
      "votes": null
    },
    {
      "id": "1146929",
      "postDate": "01/10/2021 07:15:34",
      "content": "<blockquote>\n  <p>I think LSTM is better at capturing nearby information than Transformer, but Transformer is better at capturing distant information. That's why I included both.</p>\n</blockquote>\n<p>Thank you for the explanation. I understand you try to have each layers capture different information.<br>\nI will try by myself, thanks!</p>\n<blockquote>\n  <p>I included cumulative and average features for all previous problems of the user, and features that greatly increased the score of LightGBM.</p>\n</blockquote>\n<p>Sorry for my unclear asking. I meant your NN seems to have two branches one gets features like content id, lag_time..(left side), the another one gets features like maybe cumulative and average predictions you mentioned (right side). How did you split features into two group?</p>",
      "rawMarkdown": "> I think LSTM is better at capturing nearby information than Transformer, but Transformer is better at capturing distant information. That's why I included both.\n\nThank you for the explanation. I understand you try to have each layers capture different information.\nI will try by myself, thanks!\n\n> I included cumulative and average features for all previous problems of the user, and features that greatly increased the score of LightGBM.\n\nSorry for my unclear asking. I meant your NN seems to have two branches one gets features like content id, lag_time..(left side), the another one gets features like maybe cumulative and average predictions you mentioned (right side). How did you split features into two group?",
      "votes": null
    },
    {
      "id": "1146941",
      "postDate": "01/10/2021 07:35:38",
      "content": "<p>I'm sorry, I misunderstood.<br>\nThe sequence (content_id, answered correctly etc…) are on the Transformer side, and the other features that are not sequences as explained earlier are on the Dence side. Is that the answer?</p>",
      "rawMarkdown": "I'm sorry, I misunderstood.\nThe sequence (content_id, answered correctly etc...) are on the Transformer side, and the other features that are not sequences as explained earlier are on the Dence side. Is that the answer?",
      "votes": null
    },
    {
      "id": "1146959",
      "postDate": "01/10/2021 08:00:13",
      "content": "<p>Thank you for answering. I thought cumulative features seem kind of sequence features but that is wrong?</p>",
      "rawMarkdown": "Thank you for answering. I thought cumulative features seem kind of sequence features but that is wrong?",
      "votes": null
    },
    {
      "id": "1147012",
      "postDate": "01/10/2021 08:43:19",
      "content": "<p>The cumulative features can be represented to some extent by sequences, but as shown in the figure, the length of my sequences are 100, so it do not represent the entire history of the user. That's why I added the cumulative sum and average features to the dense layer.</p>",
      "rawMarkdown": "The cumulative features can be represented to some extent by sequences, but as shown in the figure, the length of my sequences are 100, so it do not represent the entire history of the user. That's why I added the cumulative sum and average features to the dense layer.",
      "votes": null
    },
    {
      "id": "1147049",
      "postDate": "01/10/2021 08:54:33",
      "content": "<p>I understand well. I will try. <br>\nThank you very much for kindly explaining and taking your precious time!</p>",
      "rawMarkdown": "I understand well. I will try. \nThank you very much for kindly explaining and taking your precious time!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1143785,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "01/08/2021 04:15:57",
      "content": "<p>Congrats on 13th place <a href=\"https://www.kaggle.com/kurupical\" target=\"_blank\">@kurupical</a> and <a href=\"https://www.kaggle.com/takoihiraokazu\" target=\"_blank\">@takoihiraokazu</a> and thanks for sharing your solution</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1144389,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "01/08/2021 12:39:23",
      "content": "<p>Very nice write-up, thanks for sharing ! And congratz on the gold finish!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1145474,
      "author_name": "zjjszj2",
      "author_url": "",
      "post_date": "01/09/2021 07:04:38",
      "content": "<p>congratulations！</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1145485,
      "author_name": "adityaecdrid",
      "author_url": "",
      "post_date": "01/09/2021 07:13:36",
      "content": "<p>Congratulations on a cool finish!</p>\n<ul>\n<li>What's same_question_index and target?</li>\n<li>distance between the same content_id -&gt; so this is the difference b/w the sequence no when a particular content_id was last seen Vs their current sequence number assigned, right?</li>\n<li>decayed features (average of answered correctly) -&gt; How did you decay them? Did you do something like <code>alpha*new_value + (1-alpha)*old_value</code>?</li>\n<li>LIT structure in EncoderLayer. What's LIT?</li>\n</ul>\n<p>Thanks a lot!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1145614,
          "author_name": "takoihiraokazu",
          "author_url": "",
          "post_date": "01/09/2021 08:57:14",
          "content": "<p>Thank you!<br>\nI answer question about my side.</p>\n<pre><code>distance between the same content_id -&gt; so this is the difference b/w the sequence no when a particular content_id was last seen Vs their current sequence number assigned, right?\n</code></pre>\n<p>Yes.</p>\n<pre><code>decayed features (average of answered correctly) -&gt; How did you decay them? Did you do something like alpha*new_value + (1-alpha)*old_value?\n</code></pre>\n<p>I decayed the value as <code>0.3 * new_vlaue + (1 - 0.3) * old_value</code> .</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1145626,
          "author_name": "kurupical",
          "author_url": "",
          "post_date": "01/09/2021 09:07:23",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> !</p>\n<blockquote>\n  <p>What's same_question_index and target?</p>\n</blockquote>\n<p>-&gt; same as takoi's distance between the same content_id</p>\n<blockquote>\n  <p>LIT structure in EncoderLayer. What's LIT?</p>\n</blockquote>\n<p>-&gt; Sorry, I forget to explain detail. <a href=\"https://arxiv.org/pdf/2012.14164.pdf\" target=\"_blank\">https://arxiv.org/pdf/2012.14164.pdf</a>. My implements is below. TransformerEncoder is copy from torch.nn.TransformerEncoder, I just replaced the linear part of src2 with a LITLayer.</p>\n<pre><code>class LITLayer(nn.Module):\n    \"\"\"\n    https://arxiv.org/pdf/2012.14164.pdf\n    \"\"\"\n    def __init__(self, input_dim, embed_dim, dropout, activation):\n        super(LITLayer, self).__init__()\n        self.input_dim = input_dim\n        self.embed_dim = embed_dim\n        self.activation = activation\n\n        self.linear1 = nn.Linear(input_dim, embed_dim)\n        self.lstm = nn.LSTM(embed_dim, embed_dim)\n        self.linear2 = nn.Linear(embed_dim, input_dim)\n\n        self.norm_lstm = nn.LayerNorm(embed_dim)\n\n        self.dropout = nn.Dropout(dropout)\n\n    def forward(self, x):\n        x = self.linear1(x)\n        x = self.activation(x)\n        x = self.dropout(x)\n        x, _ = self.lstm(x)\n        x = self.norm_lstm(x)\n        x = self.dropout(x)\n        x = self.linear2(x)\n\n        return x\n\nclass TransformerEncoderLayer(nn.Module):\n    r\"\"\"TransformerEncoderLayer is made up of self-attn and feedforward network.\n    This standard encoder layer is based on the paper \"Attention Is All You Need\".\n    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez,\n    Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in\n    Neural Information Processing Systems, pages 6000-6010. Users may modify or implement\n    in a different way during application.\n\n    Args:\n        d_model: the number of expected features in the input (required).\n        nhead: the number of heads in the multiheadattention models (required).\n        dim_feedforward: the dimension of the feedforward network model (default=2048).\n        dropout: the dropout value (default=0.1).\n        activation: the activation function of intermediate layer, relu or gelu (default=relu).\n\n    Examples::\n        &gt;&gt;&gt; encoder_layer = nn.TransformerEncoderLayer(d_model=512, nhead=8)\n        &gt;&gt;&gt; src = torch.rand(10, 32, 512)\n        &gt;&gt;&gt; out = encoder_layer(src)\n    \"\"\"\n\n    def __init__(self, d_model, nhead, dim_feedforward=256, dropout=0.1, activation=\"relu\"):\n        super(TransformerEncoderLayer, self).__init__()\n        self.self_attn = nn.MultiheadAttention(d_model, nhead, dropout=dropout)\n        # Implementation of Feedforward model\n\n        self.norm1 = nn.LayerNorm(d_model)\n        self.norm2 = nn.LayerNorm(d_model)\n        self.dropout1 = nn.Dropout(dropout)\n        self.dropout2 = nn.Dropout(dropout)\n\n        self.activation = _get_activation_fn(activation)\n\n        self.lit_layer = LITLayer(input_dim=d_model, embed_dim=dim_feedforward, dropout=dropout, activation=self.activation)\n\n    def __setstate__(self, state):\n        if 'activation' not in state:\n            state['activation'] = nn.ReLU\n        super(TransformerEncoderLayer, self).__setstate__(state)\n\n    def forward(self, src: torch.Tensor, src_mask: Optional[torch.Tensor] = None, src_key_padding_mask: Optional[torch.Tensor] = None) -&gt; torch.Tensor:\n        r\"\"\"Pass the input through the encoder layer.\n\n        Args:\n            src: the sequence to the encoder layer (required).\n            src_mask: the mask for the src sequence (optional).\n            src_key_padding_mask: the mask for the src keys per batch (optional).\n\n        Shape:\n            see the docs in Transformer class.\n        \"\"\"\n        src2 = self.self_attn(src, src, src, attn_mask=src_mask,\n                              key_padding_mask=src_key_padding_mask)[0]\n        src = src + self.dropout1(src2)\n        src = self.norm1(src)\n       # src2 = self.linear2(self.dropout(self.activation(self.linear1(src))))\n        src2 = self.lit_layer(src)\n        src = src + self.dropout2(src2)\n        src = self.norm2(src)\n        return src\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1145637,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "01/09/2021 09:14:36",
          "content": "<p>Thanks guys for the quick response!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1146884,
      "author_name": "yaaku123",
      "author_url": "",
      "post_date": "01/10/2021 06:23:05",
      "content": "<p>Thank you for sharing a great solution and congratulations on gold medal!</p>\n<p>It might be a dumb question but what makes you stack a LSTM layer on a Transformer layer?<br>\nI am studying NN and it would be a great help if you guide me to understand how LSTM is superior to Transformer.</p>\n<p>And, how did you decide which features put in transformer or dense layer? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1146914,
          "author_name": "takoihiraokazu",
          "author_url": "",
          "post_date": "01/10/2021 06:52:20",
          "content": "<p>Thanks!</p>\n<pre><code>It might be a dumb question but what makes you stack a LSTM layer on a Transformer layer?\nI am studying NN and it would be a great help if you guide me to understand how LSTM is superior to Transformer.\n</code></pre>\n<p>I think LSTM is better at capturing nearby information than Transformer, but Transformer is better at capturing distant information. That's why I included both.</p>\n<pre><code>And, how did you decide which features put in transformer or dense layer?\n</code></pre>\n<p>I included cumulative and average features for all previous problems of the user, and features that greatly increased the score of LightGBM.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1146929,
          "author_name": "yaaku123",
          "author_url": "",
          "post_date": "01/10/2021 07:15:34",
          "content": "<blockquote>\n  <p>I think LSTM is better at capturing nearby information than Transformer, but Transformer is better at capturing distant information. That's why I included both.</p>\n</blockquote>\n<p>Thank you for the explanation. I understand you try to have each layers capture different information.<br>\nI will try by myself, thanks!</p>\n<blockquote>\n  <p>I included cumulative and average features for all previous problems of the user, and features that greatly increased the score of LightGBM.</p>\n</blockquote>\n<p>Sorry for my unclear asking. I meant your NN seems to have two branches one gets features like content id, lag_time..(left side), the another one gets features like maybe cumulative and average predictions you mentioned (right side). How did you split features into two group?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1146941,
          "author_name": "takoihiraokazu",
          "author_url": "",
          "post_date": "01/10/2021 07:35:38",
          "content": "<p>I'm sorry, I misunderstood.<br>\nThe sequence (content_id, answered correctly etc…) are on the Transformer side, and the other features that are not sequences as explained earlier are on the Dence side. Is that the answer?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1146959,
          "author_name": "yaaku123",
          "author_url": "",
          "post_date": "01/10/2021 08:00:13",
          "content": "<p>Thank you for answering. I thought cumulative features seem kind of sequence features but that is wrong?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1147012,
          "author_name": "takoihiraokazu",
          "author_url": "",
          "post_date": "01/10/2021 08:43:19",
          "content": "<p>The cumulative features can be represented to some extent by sequences, but as shown in the figure, the length of my sequences are 100, so it do not represent the entire history of the user. That's why I added the cumulative sum and average features to the dense layer.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1147049,
          "author_name": "yaaku123",
          "author_url": "",
          "post_date": "01/10/2021 08:54:33",
          "content": "<p>I understand well. I will try. <br>\nThank you very much for kindly explaining and taking your precious time!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1143748": "First of all, thanks to Hoon Pyo (Tim) Jeon and Kaggle team for such an interesting competition.\nAnd congratulates to all the winning teams!\n\nThe following is the team takoi + kurupical solution.\n\n<br>\n<br>\n\n# Team takoi + kurupical Overview\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1146523%2F3f2943b27feb7d884472c201cd032d5b%2F1.jpg?generation=1610075778291231&alt=media)\n\n# validation\n@tito's validation strategy. https://www.kaggle.com/its7171/cv-strategy\n\n<br>\n<br>\n# **kurupical side**\nI would like to have three kaggler to thank.\n@takoi for inviting me to form a team. If it weren't for you, I couldn't reach this rank!\n@limerobot for sharing DSB 3rd solution. I'm beginner in transformer for time-series data, so I learned a lot from your solution!\n@wangsg for sharing notebook https://www.kaggle.com/wangsg/a-self-attentive-model-for-knowledge-tracing.!\nI used this notebook as a baseline and finally get 0.809 CV for single transformer.\n\n## model\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1146523%2F411c8ee779411bb443284b572980ef21%2F2.jpg?generation=1610075796425059&alt=media)\n\n## hyper parameters\n* 20epochs\n* AdamW(lr=1e-3, weight_decay=0.1)\n* linear_with_warmup(lr=1e-3, warmup_epoch=2)\n\n## worked for me\n* baseline (SAKT, https://www.kaggle.com/wangsg/a-self-attentive-model-for-knowledge-tracing)\n* use all data (this notebook use only last 100 history per user)\n* embedding concat (not add) and Linear layer after cat embedding(@limerobot DSB2019 3rd solution) (+0.03) \n* Add min(timestamp_delta//1000, 300) (+0.02)\n* Add \"index that user answered same content_id at last\" (+0.005)\n* Transformer Encoder n_layers 2 -> 4 (+0.002)\n* weight_decay 0.01 -> 0.1 (+0.002)\n* LIT structure in EncoderLayer (+0.002)\n\n## not worked for me\nI did over 300 experiments, and only about 20 of them were successful.\n\n* SAINT structure (Transformer Encoder/Decoder)\n* Positional Encoding\n* Consider timeseries\n  * timedelta.cumsum() / timedelta.sum()\n  * np.log10(timedelta.cumsum()).astype(int) as category feature and embedding\n  etc...\n* optimizer AdaBelief, LookAhead(Adam), RAdam\n* more n_layers(4 => 6), more embedding_dimention (256 => 512)\n* output only the end of the sequence\n* large binning for elapsed_time/timedelta (500, 1000, etc...)\n* treat elapsed_time and timedelta as continuous \n\n<br>\n<br>\n\n# **takoi side**\nI made 1 LightGBM and 8 NN models. The model that combined Transformer and LSTM had the best CV. Here is architecture and brief description.\n​\n## Transformer + LSTM\n​\n</br>\n​\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1146523%2Fcab7462101930cf01ccbdd867e3b8a68%2F2%20(1).jpeg?generation=1610097396888437&alt=media)\n\n### features\nI used 17 features. 15 features were computed per user_id. 2 features were computed per content_id.\n#### main features\n - sum of answered correctly\n - average of answered correctly\n - sum of answered correctly for tag_user_id\n - average of answered correctly for tag_user_id\n - lag time\n - lag time of same content_id\n - previous answered correctly for the same content_id\n - distance between the same content_id\n - average of answered correctly for each content_id\n - average of lag time for each content_id\n​\n</br>\n​\n## LightGBM\nI used 97 features. The following are the main features.\n- sum of answered correctly\n- average of answered correctly\n- sum of answered correctly for tag_user_id\n- average of answered correctly for tag_user_id\n- lag time\n- lag time of same part\n- lag time of same content_id\n- previous answered correctly for the same content_id\n- distance between the same content_id\n- Word2Vec features of content_id\n- decayed features (average of answered correctly)\n- number of consecutive times with the same user answer\n- flag if the user answer being answered in succession matches the correct answer\n- average of answered correctly for each content_id\n- average of lag time for each content_id",
    "1143785": "Congrats on 13th place @kurupical and @takoihiraokazu and thanks for sharing your solution",
    "1144389": "Very nice write-up, thanks for sharing ! And congratz on the gold finish!",
    "1145474": "congratulations！",
    "1145485": "Congratulations on a cool finish!\n\n- What's same_question_index and target?\n- distance between the same content_id -> so this is the difference b/w the sequence no when a particular content_id was last seen Vs their current sequence number assigned, right?\n- decayed features (average of answered correctly) -> How did you decay them? Did you do something like `alpha*new_value + (1-alpha)*old_value`?\n- LIT structure in EncoderLayer. What's LIT?\n\nThanks a lot!",
    "1145614": "Thank you!\nI answer question about my side.\n\n``` \ndistance between the same content_id -> so this is the difference b/w the sequence no when a particular content_id was last seen Vs their current sequence number assigned, right?\n```\nYes.\n\n```\ndecayed features (average of answered correctly) -> How did you decay them? Did you do something like alpha*new_value + (1-alpha)*old_value?\n```\nI decayed the value as ``` 0.3 * new_vlaue + (1 - 0.3) * old_value``` .",
    "1145626": "Hi @adityaecdrid !\n\n> What's same_question_index and target?\n\n-> same as takoi's distance between the same content_id\n\n> LIT structure in EncoderLayer. What's LIT?\n\n-> Sorry, I forget to explain detail. https://arxiv.org/pdf/2012.14164.pdf. My implements is below. TransformerEncoder is copy from torch.nn.TransformerEncoder, I just replaced the linear part of src2 with a LITLayer.\n\n```\nclass LITLayer(nn.Module):\n    \"\"\"\n    https://arxiv.org/pdf/2012.14164.pdf\n    \"\"\"\n    def __init__(self, input_dim, embed_dim, dropout, activation):\n        super(LITLayer, self).__init__()\n        self.input_dim = input_dim\n        self.embed_dim = embed_dim\n        self.activation = activation\n\n        self.linear1 = nn.Linear(input_dim, embed_dim)\n        self.lstm = nn.LSTM(embed_dim, embed_dim)\n        self.linear2 = nn.Linear(embed_dim, input_dim)\n\n        self.norm_lstm = nn.LayerNorm(embed_dim)\n\n        self.dropout = nn.Dropout(dropout)\n\n    def forward(self, x):\n        x = self.linear1(x)\n        x = self.activation(x)\n        x = self.dropout(x)\n        x, _ = self.lstm(x)\n        x = self.norm_lstm(x)\n        x = self.dropout(x)\n        x = self.linear2(x)\n\n        return x\n\nclass TransformerEncoderLayer(nn.Module):\n    r\"\"\"TransformerEncoderLayer is made up of self-attn and feedforward network.\n    This standard encoder layer is based on the paper \"Attention Is All You Need\".\n    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez,\n    Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in\n    Neural Information Processing Systems, pages 6000-6010. Users may modify or implement\n    in a different way during application.\n\n    Args:\n        d_model: the number of expected features in the input (required).\n        nhead: the number of heads in the multiheadattention models (required).\n        dim_feedforward: the dimension of the feedforward network model (default=2048).\n        dropout: the dropout value (default=0.1).\n        activation: the activation function of intermediate layer, relu or gelu (default=relu).\n\n    Examples::\n        >>> encoder_layer = nn.TransformerEncoderLayer(d_model=512, nhead=8)\n        >>> src = torch.rand(10, 32, 512)\n        >>> out = encoder_layer(src)\n    \"\"\"\n\n    def __init__(self, d_model, nhead, dim_feedforward=256, dropout=0.1, activation=\"relu\"):\n        super(TransformerEncoderLayer, self).__init__()\n        self.self_attn = nn.MultiheadAttention(d_model, nhead, dropout=dropout)\n        # Implementation of Feedforward model\n\n        self.norm1 = nn.LayerNorm(d_model)\n        self.norm2 = nn.LayerNorm(d_model)\n        self.dropout1 = nn.Dropout(dropout)\n        self.dropout2 = nn.Dropout(dropout)\n\n        self.activation = _get_activation_fn(activation)\n\n        self.lit_layer = LITLayer(input_dim=d_model, embed_dim=dim_feedforward, dropout=dropout, activation=self.activation)\n\n    def __setstate__(self, state):\n        if 'activation' not in state:\n            state['activation'] = nn.ReLU\n        super(TransformerEncoderLayer, self).__setstate__(state)\n\n    def forward(self, src: torch.Tensor, src_mask: Optional[torch.Tensor] = None, src_key_padding_mask: Optional[torch.Tensor] = None) -> torch.Tensor:\n        r\"\"\"Pass the input through the encoder layer.\n\n        Args:\n            src: the sequence to the encoder layer (required).\n            src_mask: the mask for the src sequence (optional).\n            src_key_padding_mask: the mask for the src keys per batch (optional).\n\n        Shape:\n            see the docs in Transformer class.\n        \"\"\"\n        src2 = self.self_attn(src, src, src, attn_mask=src_mask,\n                              key_padding_mask=src_key_padding_mask)[0]\n        src = src + self.dropout1(src2)\n        src = self.norm1(src)\n       # src2 = self.linear2(self.dropout(self.activation(self.linear1(src))))\n        src2 = self.lit_layer(src)\n        src = src + self.dropout2(src2)\n        src = self.norm2(src)\n        return src\n```",
    "1145637": "Thanks guys for the quick response!",
    "1146884": "Thank you for sharing a great solution and congratulations on gold medal!\n\nIt might be a dumb question but what makes you stack a LSTM layer on a Transformer layer?\nI am studying NN and it would be a great help if you guide me to understand how LSTM is superior to Transformer.\n\nAnd, how did you decide which features put in transformer or dense layer?",
    "1146914": "Thanks!\n```\nIt might be a dumb question but what makes you stack a LSTM layer on a Transformer layer?\nI am studying NN and it would be a great help if you guide me to understand how LSTM is superior to Transformer.\n```\nI think LSTM is better at capturing nearby information than Transformer, but Transformer is better at capturing distant information. That's why I included both.\n\n```\nAnd, how did you decide which features put in transformer or dense layer?\n```\nI included cumulative and average features for all previous problems of the user, and features that greatly increased the score of LightGBM.",
    "1146929": "> I think LSTM is better at capturing nearby information than Transformer, but Transformer is better at capturing distant information. That's why I included both.\n\nThank you for the explanation. I understand you try to have each layers capture different information.\nI will try by myself, thanks!\n\n> I included cumulative and average features for all previous problems of the user, and features that greatly increased the score of LightGBM.\n\nSorry for my unclear asking. I meant your NN seems to have two branches one gets features like content id, lag_time..(left side), the another one gets features like maybe cumulative and average predictions you mentioned (right side). How did you split features into two group?",
    "1146941": "I'm sorry, I misunderstood.\nThe sequence (content_id, answered correctly etc...) are on the Transformer side, and the other features that are not sequences as explained earlier are on the Dence side. Is that the answer?",
    "1146959": "Thank you for answering. I thought cumulative features seem kind of sequence features but that is wrong?",
    "1147012": "The cumulative features can be represented to some extent by sequences, but as shown in the figure, the length of my sequences are 100, so it do not represent the entire history of the user. That's why I added the cumulative sum and average features to the dense layer.",
    "1147049": "I understand well. I will try. \nThank you very much for kindly explaining and taking your precious time!"
  },
  "source": "meta"
}