{
  "id": 209713,
  "title": "17th place solution - 4 features model",
  "url": "/competitions/riiid-test-answer-prediction/writeups/saintnikola-17th-place-solution-4-features-model",
  "author_name": "",
  "post_date": "2021-02-19T13:19:48.507Z",
  "votes": 63,
  "comment_count": 20,
  "views": 0,
  "content": "<p>My solution is based on a stack of the 3 models:</p>\n<p>Edit: <a href=\"https://github.com/NikolaBacic/riiid\" target=\"_blank\">Code</a></p>\n<p>Transformer(Encoder) - Validation: 0.8100<br>\nLSTM - Validation: 0.8062<br>\nGRU - Validation: 0.8060</p>\n<p>LightGBM for stack: Validation: 0.8119 LB: 0.814x</p>\n<p>I validated on new users (2.5M rows). <a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">This one</a> is great, but it was computationally expensive for me. </p>\n<p>All three models used 4 features (embeddings):</p>\n<ul>\n<li>question id embedding</li>\n<li>response of the previous question (1-correct 0-incorrect)</li>\n<li><strong>ln(lag+1) * minute_embedding : taking a ln(x+1) of the lag initialy improved my score by ~0.01. I'd be very interested to hear whether it'd improve your models too</strong></li>\n<li>ln(prior_elapsed_time+1) * minute_embedding</li>\n</ul>\n<p>Input is sum of these 4 embeddings.</p>\n<p><strong>Parameters</strong></p>\n<p>Shareable parameters (all three models):<br>\nmax_quest = 300 (window size)<br>\nslide = 150<br>\nAdam optimizer<br>\ncosine lr scheduler<br>\nBCE loss<br>\n<strong>xavier_uniform_ weight initialization (0.004 improvement over PyTorch's default one)</strong></p>\n<p>I used Optuna for hyperparameter search on 20% of the data. It was my first time using it, and it's a great tool!</p>\n<p>Transformer hyperparameters:<br>\nnhead = 8<br>\nhead_dim = 60<br>\ndim_feedforward = 2048<br>\nnum_encoder_layers = 8<br>\nepochs = 6<br>\nbatch_size = 64<br>\nlr = 0.00019809259513409007<br>\nwarmup_steps = 150*5</p>\n<p>LSTM hyperparameters:<br>\ninput_size_lstm = 384<br>\nhidden_size_lstm = 768<br>\nnum_layers_lstm = 4<br>\nepochs = 4<br>\nbatch_size = 64<br>\nlr = 0.0007019926812886481<br>\nwarmup_steps = 100</p>\n<p>GRU hyperparameters:<br>\ninput_size_gru = 320<br>\nhidden_size_gru = 512<br>\nnum_layers_gru = 3<br>\nepochs = 4<br>\nbatch_size = 64<br>\nlr = 0.0008419253431185227<br>\nwarmup_steps = 80</p>\n<p>What didn't work:<br>\n-lectures<br>\n-questions metadata<br>\n-position encoding: both learnable and hard-coded<br>\n-predicting user_answer instead of correctness<br>\n-gradient clipping<br>\n-label smoothing<br>\n-dropout<br>\n-…</p>\n<p>gg</p>",
  "messages": [
    {
      "id": "1144241",
      "postDate": "01/08/2021 10:44:14",
      "content": "<p>My solution is based on a stack of the 3 models:</p>\n<p>Edit: <a href=\"https://github.com/NikolaBacic/riiid\" target=\"_blank\">Code</a></p>\n<p>Transformer(Encoder) - Validation: 0.8100<br>\nLSTM - Validation: 0.8062<br>\nGRU - Validation: 0.8060</p>\n<p>LightGBM for stack: Validation: 0.8119 LB: 0.814x</p>\n<p>I validated on new users (2.5M rows). <a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">This one</a> is great, but it was computationally expensive for me. </p>\n<p>All three models used 4 features (embeddings):</p>\n<ul>\n<li>question id embedding</li>\n<li>response of the previous question (1-correct 0-incorrect)</li>\n<li><strong>ln(lag+1) * minute_embedding : taking a ln(x+1) of the lag initialy improved my score by ~0.01. I'd be very interested to hear whether it'd improve your models too</strong></li>\n<li>ln(prior_elapsed_time+1) * minute_embedding</li>\n</ul>\n<p>Input is sum of these 4 embeddings.</p>\n<p><strong>Parameters</strong></p>\n<p>Shareable parameters (all three models):<br>\nmax_quest = 300 (window size)<br>\nslide = 150<br>\nAdam optimizer<br>\ncosine lr scheduler<br>\nBCE loss<br>\n<strong>xavier_uniform_ weight initialization (0.004 improvement over PyTorch's default one)</strong></p>\n<p>I used Optuna for hyperparameter search on 20% of the data. It was my first time using it, and it's a great tool!</p>\n<p>Transformer hyperparameters:<br>\nnhead = 8<br>\nhead_dim = 60<br>\ndim_feedforward = 2048<br>\nnum_encoder_layers = 8<br>\nepochs = 6<br>\nbatch_size = 64<br>\nlr = 0.00019809259513409007<br>\nwarmup_steps = 150*5</p>\n<p>LSTM hyperparameters:<br>\ninput_size_lstm = 384<br>\nhidden_size_lstm = 768<br>\nnum_layers_lstm = 4<br>\nepochs = 4<br>\nbatch_size = 64<br>\nlr = 0.0007019926812886481<br>\nwarmup_steps = 100</p>\n<p>GRU hyperparameters:<br>\ninput_size_gru = 320<br>\nhidden_size_gru = 512<br>\nnum_layers_gru = 3<br>\nepochs = 4<br>\nbatch_size = 64<br>\nlr = 0.0008419253431185227<br>\nwarmup_steps = 80</p>\n<p>What didn't work:<br>\n-lectures<br>\n-questions metadata<br>\n-position encoding: both learnable and hard-coded<br>\n-predicting user_answer instead of correctness<br>\n-gradient clipping<br>\n-label smoothing<br>\n-dropout<br>\n-…</p>\n<p>gg</p>",
      "rawMarkdown": "My solution is based on a stack of the 3 models:\n\nEdit: [Code](https://github.com/NikolaBacic/riiid)\n\nTransformer(Encoder) - Validation: 0.8100\nLSTM - Validation: 0.8062\nGRU - Validation: 0.8060\n\nLightGBM for stack: Validation: 0.8119 LB: 0.814x\n\nI validated on new users (2.5M rows). [This one](https://www.kaggle.com/its7171/cv-strategy) is great, but it was computationally expensive for me. \n\nAll three models used 4 features (embeddings):\n- question id embedding\n- response of the previous question (1-correct 0-incorrect)\n- **ln(lag+1) * minute_embedding : taking a ln(x+1) of the lag initialy improved my score by ~0.01. I'd be very interested to hear whether it'd improve your models too**\n- ln(prior_elapsed_time+1) * minute_embedding\n\nInput is sum of these 4 embeddings.\n\n**Parameters**\n\nShareable parameters (all three models):\nmax_quest = 300 (window size)\nslide = 150\nAdam optimizer\ncosine lr scheduler\nBCE loss\n**xavier_uniform_ weight initialization (0.004 improvement over PyTorch's default one)**\n\nI used Optuna for hyperparameter search on 20% of the data. It was my first time using it, and it's a great tool!\n\nTransformer hyperparameters:\nnhead = 8\nhead_dim = 60\ndim_feedforward = 2048\nnum_encoder_layers = 8\nepochs = 6\nbatch_size = 64\nlr = 0.00019809259513409007\nwarmup_steps = 150*5\n\nLSTM hyperparameters:\ninput_size_lstm = 384\nhidden_size_lstm = 768\nnum_layers_lstm = 4\nepochs = 4\nbatch_size = 64\nlr = 0.0007019926812886481\nwarmup_steps = 100\n\nGRU hyperparameters:\ninput_size_gru = 320\nhidden_size_gru = 512\nnum_layers_gru = 3\nepochs = 4\nbatch_size = 64\nlr = 0.0008419253431185227\nwarmup_steps = 80\n\nWhat didn't work:\n-lectures\n-questions metadata\n-position encoding: both learnable and hard-coded\n-predicting user_answer instead of correctness\n-gradient clipping\n-label smoothing\n-dropout\n-...\n\ngg",
      "votes": null
    },
    {
      "id": "1144253",
      "postDate": "01/08/2021 10:53:27",
      "content": "<p>Cool, thanks for sharing!</p>",
      "rawMarkdown": "Cool, thanks for sharing!",
      "votes": null
    },
    {
      "id": "1144256",
      "postDate": "01/08/2021 10:55:02",
      "content": "<p>Congrats !! Only 4 features is pretty amazing<br>\n-Response of the previous question is for quesiton_id or for task_container_id ?<br>\n-Did you try to sum mutiple previous responses ? Like a rolling window</p>",
      "rawMarkdown": "Congrats !! Only 4 features is pretty amazing\n-Response of the previous question is for quesiton_id or for task_container_id ?\n-Did you try to sum mutiple previous responses ? Like a rolling window",
      "votes": null
    },
    {
      "id": "1144263",
      "postDate": "01/08/2021 10:59:08",
      "content": "<p>Part and tags improved my score on 10% of the data, but with 100% I found that transformer learns them :)<br>\nThanks!</p>",
      "rawMarkdown": "Part and tags improved my score on 10% of the data, but with 100% I found that transformer learns them :)\nThanks!",
      "votes": null
    },
    {
      "id": "1144265",
      "postDate": "01/08/2021 10:59:27",
      "content": "<p>Congrats on 17th place and thanks for sharing details solution <a href=\"https://www.kaggle.com/bacicnikola\" target=\"_blank\">@bacicnikola</a> </p>",
      "rawMarkdown": "Congrats on 17th place and thanks for sharing details solution @bacicnikola",
      "votes": null
    },
    {
      "id": "1144269",
      "postDate": "01/08/2021 11:03:32",
      "content": "<blockquote>\n  <p>-Response of the previous question is for quesiton_id or for task_container_id ?</p>\n</blockquote>\n<p>It's for quesiton_id. Yes, there is some leakage, but when I tried to fix that I got some funny results. Maybe my implementation wasn't right :? </p>\n<blockquote>\n  <p>-Did you try to sum mutiple previous responses ? Like a rolling window</p>\n</blockquote>\n<p>I didn't. I think that transformer can pick that up, right?</p>",
      "rawMarkdown": "> -Response of the previous question is for quesiton_id or for task_container_id ?\n\nIt's for quesiton_id. Yes, there is some leakage, but when I tried to fix that I got some funny results. Maybe my implementation wasn't right :? \n\n> -Did you try to sum mutiple previous responses ? Like a rolling window\n\nI didn't. I think that transformer can pick that up, right?",
      "votes": null
    },
    {
      "id": "1144378",
      "postDate": "01/08/2021 12:32:46",
      "content": "<p>Simple yet effective ! Congratz on the strong finish !</p>",
      "rawMarkdown": "Simple yet effective ! Congratz on the strong finish !",
      "votes": null
    },
    {
      "id": "1144467",
      "postDate": "01/08/2021 13:30:28",
      "content": "<p>Hi, congrats! Can you please elaborate on <code>max_quest</code> and <code>slide</code>?</p>",
      "rawMarkdown": "Hi, congrats! Can you please elaborate on `max_quest` and `slide`?",
      "votes": null
    },
    {
      "id": "1144570",
      "postDate": "01/08/2021 14:40:26",
      "content": "<p>This is related to the sampling strategy. <code>max_quest</code> is maximum number of questions processed at a time i.e. maximum input sequence length.<br>\nNow, some users might have <code>&gt; max_quest</code> questions in their histories so you need to make multiple samples from one user. This is when <code>slide</code> comes into play. The simplest way to handle this is to set <code>slide = max_quest</code> and get samples from 0-300, 300-600, 600-900…<br>\nBut this is suboptimal because questions close to the left hand side of the sample, like 300:350, or 600-650 can't see recent history. So I set <code>slide = 150</code> and make samples like this:</p>\n<p>sample1: 0-300 -&gt; calculate loss on all positions<br>\nsample2: 150-450 -&gt; calculate loss on 300-450 positions<br>\nsample3: 300-600 -&gt; calculate loss on 450-600 positions<br>\netc.</p>\n<p>Hopefully this is clear enough.</p>",
      "rawMarkdown": "This is related to the sampling strategy. `max_quest` is maximum number of questions processed at a time i.e. maximum input sequence length.\nNow, some users might have ` > max_quest` questions in their histories so you need to make multiple samples from one user. This is when `slide` comes into play. The simplest way to handle this is to set `slide = max_quest` and get samples from 0-300, 300-600, 600-900...\nBut this is suboptimal because questions close to the left hand side of the sample, like 300:350, or 600-650 can't see recent history. So I set `slide = 150` and make samples like this:\n\nsample1: 0-300 -> calculate loss on all positions\nsample2: 150-450 -> calculate loss on 300-450 positions\nsample3: 300-600 -> calculate loss on 450-600 positions\netc.\n\nHopefully this is clear enough.",
      "votes": null
    },
    {
      "id": "1144952",
      "postDate": "01/08/2021 19:13:39",
      "content": "<p>Congrats and thank you for sharing!</p>",
      "rawMarkdown": "Congrats and thank you for sharing!",
      "votes": null
    },
    {
      "id": "1145221",
      "postDate": "01/09/2021 01:35:09",
      "content": "<p>Thanks a lot for the clear explanation. One more question related to <code>minute_embedding</code>, please. Is it <code>lag</code> and <code>prior_elapsed</code> converted to minutes, then multiplied by a shared learnable embedding?</p>",
      "rawMarkdown": "Thanks a lot for the clear explanation. One more question related to `minute_embedding`, please. Is it `lag` and `prior_elapsed` converted to minutes, then multiplied by a shared learnable embedding?",
      "votes": null
    },
    {
      "id": "1145329",
      "postDate": "01/09/2021 04:29:51",
      "content": "<p>thanks for sharing! a great job!</p>",
      "rawMarkdown": "thanks for sharing! a great job!",
      "votes": null
    },
    {
      "id": "1145843",
      "postDate": "01/09/2021 11:22:11",
      "content": "<p>Yes, but embedding is not shared. Not that it makes too much of a difference;  I remember trying that and it giving me negligible worse result.</p>",
      "rawMarkdown": "Yes, but embedding is not shared. Not that it makes too much of a difference;  I remember trying that and it giving me negligible worse result.",
      "votes": null
    },
    {
      "id": "1146286",
      "postDate": "01/09/2021 16:46:21",
      "content": "<p>Simple and clever nice!! Regarding the inferenece did you strugle or had to optimize the code?</p>\n<p>Edit: any tips on this are welcome</p>\n<p>Ps: I tried also last day to stack 5 weak models but didn't pass the inference stage and didn't have time to optimize it further (cv meta-model 0.805) - It was more as proof of concept and for learning rather than hitting the LB.. I should have tried with 3 though</p>",
      "rawMarkdown": "Simple and clever nice!! Regarding the inferenece did you strugle or had to optimize the code?\n\nEdit: any tips on this are welcome\n\nPs: I tried also last day to stack 5 weak models but didn't pass the inference stage and didn't have time to optimize it further (cv meta-model 0.805) - It was more as proof of concept and for learning rather than hitting the LB.. I should have tried with 3 though",
      "votes": null
    },
    {
      "id": "1146394",
      "postDate": "01/09/2021 18:36:54",
      "content": "<p>Great, it was really fun to compete with you!</p>",
      "rawMarkdown": "Great, it was really fun to compete with you!",
      "votes": null
    },
    {
      "id": "1146528",
      "postDate": "01/09/2021 20:29:02",
      "content": "<p>Ofcourse I struggled, out of first 13 submission, 4 were successful :) All three of my models preprocess the data in the same way, so I had to do it only once!</p>\n<p>I preprocessed the data locally into python dictionary with the following function, and uploaded it as a kaggle dataset. During inference, I was simply updating the inner dictionaries.</p>\n<pre><code>def csv_to_dict(df):\n    \"\"\"maps the training data from csv to dictionary\"\"\"\n    # separate question events\n    questions_df = df[df[\"content_type_id\"] == False]\n\n    # fill nans\n    questions_df[\"prior_question_elapsed_time\"] = questions_df[\"prior_question_elapsed_time\"].fillna(301000)\n\n    # questions history container\n    questions_container = dict(tuple(questions_df.groupby([\"user_id\"])))\n    # df -&gt; tensors\n    for user_id in questions_container.keys():\n        tmp_df = questions_container[user_id]\n        questions_container[user_id] = {\n            \"content_id\": torch.LongTensor(tmp_df[\"content_id\"].values),\n            \"timestamp\": torch.LongTensor(tmp_df[\"timestamp\"].values),\n            \"prior_question_elapsed_time\": torch.LongTensor(tmp_df[\"prior_question_elapsed_time\"].values),\n            \"answered_correctly\": torch.ShortTensor(tmp_df[\"answered_correctly\"].values),\n            }\n\n    return questions_container\n</code></pre>",
      "rawMarkdown": "Ofcourse I struggled, out of first 13 submission, 4 were successful :) All three of my models preprocess the data in the same way, so I had to do it only once!\n\nI preprocessed the data locally into python dictionary with the following function, and uploaded it as a kaggle dataset. During inference, I was simply updating the inner dictionaries.\n\n```\ndef csv_to_dict(df):\n    \"\"\"maps the training data from csv to dictionary\"\"\"\n    # separate question events\n    questions_df = df[df[\"content_type_id\"] == False]\n        \n    # fill nans\n    questions_df[\"prior_question_elapsed_time\"] = questions_df[\"prior_question_elapsed_time\"].fillna(301000)\n        \n    # questions history container\n    questions_container = dict(tuple(questions_df.groupby([\"user_id\"])))\n    # df -> tensors\n    for user_id in questions_container.keys():\n        tmp_df = questions_container[user_id]\n        questions_container[user_id] = {\n            \"content_id\": torch.LongTensor(tmp_df[\"content_id\"].values),\n            \"timestamp\": torch.LongTensor(tmp_df[\"timestamp\"].values),\n            \"prior_question_elapsed_time\": torch.LongTensor(tmp_df[\"prior_question_elapsed_time\"].values),\n            \"answered_correctly\": torch.ShortTensor(tmp_df[\"answered_correctly\"].values),\n            }\n            \n    return questions_container\n\n```",
      "votes": null
    },
    {
      "id": "1146533",
      "postDate": "01/09/2021 20:35:36",
      "content": "<p>Thanks, you were pushing me to limit, sorry I couldn't keep up :)  I was really tired last 2-3 weeks  (not that it's an excuse) and couldn't make almost no progress.</p>",
      "rawMarkdown": "Thanks, you were pushing me to limit, sorry I couldn't keep up :)  I was really tired last 2-3 weeks  (not that it's an excuse) and couldn't make almost no progress.",
      "votes": null
    },
    {
      "id": "1149597",
      "postDate": "01/12/2021 01:59:03",
      "content": "<p>Congrats~~~ Your work is really impressive!<br>\nI am a rookie in data science and I curiosity about how to generate 'question id embedding' features. Is it one-hot-encoding vector of question id?</p>",
      "rawMarkdown": "Congrats~~~ Your work is really impressive!\nI am a rookie in data science and I curiosity about how to generate 'question id embedding' features. Is it one-hot-encoding vector of question id?",
      "votes": null
    },
    {
      "id": "1149952",
      "postDate": "01/12/2021 09:11:43",
      "content": "<p>No, it's a vector repsesentation: <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.Embedding.html\" target=\"_blank\">Embedding</a></p>",
      "rawMarkdown": "No, it's a vector repsesentation: [Embedding](https://pytorch.org/docs/stable/generated/torch.nn.Embedding.html)",
      "votes": null
    },
    {
      "id": "1155694",
      "postDate": "01/16/2021 15:26:02",
      "content": "<p>Wait. Did you say you only use 4 features? That's impressive!</p>",
      "rawMarkdown": "Wait. Did you say you only use 4 features? That's impressive!",
      "votes": null
    },
    {
      "id": "1155965",
      "postDate": "01/16/2021 20:04:07",
      "content": "<p>Well who is learning, me or the model? :)</p>",
      "rawMarkdown": "Well who is learning, me or the model? :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1144253,
      "author_name": "chuxianmo",
      "author_url": "",
      "post_date": "01/08/2021 10:53:27",
      "content": "<p>Cool, thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1144256,
      "author_name": "alexj21",
      "author_url": "",
      "post_date": "01/08/2021 10:55:02",
      "content": "<p>Congrats !! Only 4 features is pretty amazing<br>\n-Response of the previous question is for quesiton_id or for task_container_id ?<br>\n-Did you try to sum mutiple previous responses ? Like a rolling window</p>",
      "votes": null,
      "replies": [
        {
          "id": 1144263,
          "author_name": "bacicnikola",
          "author_url": "",
          "post_date": "01/08/2021 10:59:08",
          "content": "<p>Part and tags improved my score on 10% of the data, but with 100% I found that transformer learns them :)<br>\nThanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1144269,
          "author_name": "bacicnikola",
          "author_url": "",
          "post_date": "01/08/2021 11:03:32",
          "content": "<blockquote>\n  <p>-Response of the previous question is for quesiton_id or for task_container_id ?</p>\n</blockquote>\n<p>It's for quesiton_id. Yes, there is some leakage, but when I tried to fix that I got some funny results. Maybe my implementation wasn't right :? </p>\n<blockquote>\n  <p>-Did you try to sum mutiple previous responses ? Like a rolling window</p>\n</blockquote>\n<p>I didn't. I think that transformer can pick that up, right?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1144265,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "01/08/2021 10:59:27",
      "content": "<p>Congrats on 17th place and thanks for sharing details solution <a href=\"https://www.kaggle.com/bacicnikola\" target=\"_blank\">@bacicnikola</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1144378,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "01/08/2021 12:32:46",
      "content": "<p>Simple yet effective ! Congratz on the strong finish !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1144467,
      "author_name": "yuntai",
      "author_url": "",
      "post_date": "01/08/2021 13:30:28",
      "content": "<p>Hi, congrats! Can you please elaborate on <code>max_quest</code> and <code>slide</code>?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1144570,
          "author_name": "bacicnikola",
          "author_url": "",
          "post_date": "01/08/2021 14:40:26",
          "content": "<p>This is related to the sampling strategy. <code>max_quest</code> is maximum number of questions processed at a time i.e. maximum input sequence length.<br>\nNow, some users might have <code>&gt; max_quest</code> questions in their histories so you need to make multiple samples from one user. This is when <code>slide</code> comes into play. The simplest way to handle this is to set <code>slide = max_quest</code> and get samples from 0-300, 300-600, 600-900…<br>\nBut this is suboptimal because questions close to the left hand side of the sample, like 300:350, or 600-650 can't see recent history. So I set <code>slide = 150</code> and make samples like this:</p>\n<p>sample1: 0-300 -&gt; calculate loss on all positions<br>\nsample2: 150-450 -&gt; calculate loss on 300-450 positions<br>\nsample3: 300-600 -&gt; calculate loss on 450-600 positions<br>\netc.</p>\n<p>Hopefully this is clear enough.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1145221,
          "author_name": "yuntai",
          "author_url": "",
          "post_date": "01/09/2021 01:35:09",
          "content": "<p>Thanks a lot for the clear explanation. One more question related to <code>minute_embedding</code>, please. Is it <code>lag</code> and <code>prior_elapsed</code> converted to minutes, then multiplied by a shared learnable embedding?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1145843,
          "author_name": "bacicnikola",
          "author_url": "",
          "post_date": "01/09/2021 11:22:11",
          "content": "<p>Yes, but embedding is not shared. Not that it makes too much of a difference;  I remember trying that and it giving me negligible worse result.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1144952,
      "author_name": "sishihara",
      "author_url": "",
      "post_date": "01/08/2021 19:13:39",
      "content": "<p>Congrats and thank you for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1145329,
      "author_name": "zjjszj2",
      "author_url": "",
      "post_date": "01/09/2021 04:29:51",
      "content": "<p>thanks for sharing! a great job!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1146286,
      "author_name": "imeintanis",
      "author_url": "",
      "post_date": "01/09/2021 16:46:21",
      "content": "<p>Simple and clever nice!! Regarding the inferenece did you strugle or had to optimize the code?</p>\n<p>Edit: any tips on this are welcome</p>\n<p>Ps: I tried also last day to stack 5 weak models but didn't pass the inference stage and didn't have time to optimize it further (cv meta-model 0.805) - It was more as proof of concept and for learning rather than hitting the LB.. I should have tried with 3 though</p>",
      "votes": null,
      "replies": [
        {
          "id": 1146528,
          "author_name": "bacicnikola",
          "author_url": "",
          "post_date": "01/09/2021 20:29:02",
          "content": "<p>Ofcourse I struggled, out of first 13 submission, 4 were successful :) All three of my models preprocess the data in the same way, so I had to do it only once!</p>\n<p>I preprocessed the data locally into python dictionary with the following function, and uploaded it as a kaggle dataset. During inference, I was simply updating the inner dictionaries.</p>\n<pre><code>def csv_to_dict(df):\n    \"\"\"maps the training data from csv to dictionary\"\"\"\n    # separate question events\n    questions_df = df[df[\"content_type_id\"] == False]\n\n    # fill nans\n    questions_df[\"prior_question_elapsed_time\"] = questions_df[\"prior_question_elapsed_time\"].fillna(301000)\n\n    # questions history container\n    questions_container = dict(tuple(questions_df.groupby([\"user_id\"])))\n    # df -&gt; tensors\n    for user_id in questions_container.keys():\n        tmp_df = questions_container[user_id]\n        questions_container[user_id] = {\n            \"content_id\": torch.LongTensor(tmp_df[\"content_id\"].values),\n            \"timestamp\": torch.LongTensor(tmp_df[\"timestamp\"].values),\n            \"prior_question_elapsed_time\": torch.LongTensor(tmp_df[\"prior_question_elapsed_time\"].values),\n            \"answered_correctly\": torch.ShortTensor(tmp_df[\"answered_correctly\"].values),\n            }\n\n    return questions_container\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1146394,
      "author_name": "mamasinkgs",
      "author_url": "",
      "post_date": "01/09/2021 18:36:54",
      "content": "<p>Great, it was really fun to compete with you!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1146533,
          "author_name": "bacicnikola",
          "author_url": "",
          "post_date": "01/09/2021 20:35:36",
          "content": "<p>Thanks, you were pushing me to limit, sorry I couldn't keep up :)  I was really tired last 2-3 weeks  (not that it's an excuse) and couldn't make almost no progress.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1149597,
      "author_name": "tonyleekuangchen",
      "author_url": "",
      "post_date": "01/12/2021 01:59:03",
      "content": "<p>Congrats~~~ Your work is really impressive!<br>\nI am a rookie in data science and I curiosity about how to generate 'question id embedding' features. Is it one-hot-encoding vector of question id?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1149952,
          "author_name": "bacicnikola",
          "author_url": "",
          "post_date": "01/12/2021 09:11:43",
          "content": "<p>No, it's a vector repsesentation: <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.Embedding.html\" target=\"_blank\">Embedding</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1155694,
      "author_name": "faraksuli",
      "author_url": "",
      "post_date": "01/16/2021 15:26:02",
      "content": "<p>Wait. Did you say you only use 4 features? That's impressive!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1155965,
          "author_name": "bacicnikola",
          "author_url": "",
          "post_date": "01/16/2021 20:04:07",
          "content": "<p>Well who is learning, me or the model? :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1144241": "My solution is based on a stack of the 3 models:\n\nEdit: [Code](https://github.com/NikolaBacic/riiid)\n\nTransformer(Encoder) - Validation: 0.8100\nLSTM - Validation: 0.8062\nGRU - Validation: 0.8060\n\nLightGBM for stack: Validation: 0.8119 LB: 0.814x\n\nI validated on new users (2.5M rows). [This one](https://www.kaggle.com/its7171/cv-strategy) is great, but it was computationally expensive for me. \n\nAll three models used 4 features (embeddings):\n- question id embedding\n- response of the previous question (1-correct 0-incorrect)\n- **ln(lag+1) * minute_embedding : taking a ln(x+1) of the lag initialy improved my score by ~0.01. I'd be very interested to hear whether it'd improve your models too**\n- ln(prior_elapsed_time+1) * minute_embedding\n\nInput is sum of these 4 embeddings.\n\n**Parameters**\n\nShareable parameters (all three models):\nmax_quest = 300 (window size)\nslide = 150\nAdam optimizer\ncosine lr scheduler\nBCE loss\n**xavier_uniform_ weight initialization (0.004 improvement over PyTorch's default one)**\n\nI used Optuna for hyperparameter search on 20% of the data. It was my first time using it, and it's a great tool!\n\nTransformer hyperparameters:\nnhead = 8\nhead_dim = 60\ndim_feedforward = 2048\nnum_encoder_layers = 8\nepochs = 6\nbatch_size = 64\nlr = 0.00019809259513409007\nwarmup_steps = 150*5\n\nLSTM hyperparameters:\ninput_size_lstm = 384\nhidden_size_lstm = 768\nnum_layers_lstm = 4\nepochs = 4\nbatch_size = 64\nlr = 0.0007019926812886481\nwarmup_steps = 100\n\nGRU hyperparameters:\ninput_size_gru = 320\nhidden_size_gru = 512\nnum_layers_gru = 3\nepochs = 4\nbatch_size = 64\nlr = 0.0008419253431185227\nwarmup_steps = 80\n\nWhat didn't work:\n-lectures\n-questions metadata\n-position encoding: both learnable and hard-coded\n-predicting user_answer instead of correctness\n-gradient clipping\n-label smoothing\n-dropout\n-...\n\ngg",
    "1144253": "Cool, thanks for sharing!",
    "1144256": "Congrats !! Only 4 features is pretty amazing\n-Response of the previous question is for quesiton_id or for task_container_id ?\n-Did you try to sum mutiple previous responses ? Like a rolling window",
    "1144263": "Part and tags improved my score on 10% of the data, but with 100% I found that transformer learns them :)\nThanks!",
    "1144265": "Congrats on 17th place and thanks for sharing details solution @bacicnikola",
    "1144269": "> -Response of the previous question is for quesiton_id or for task_container_id ?\n\nIt's for quesiton_id. Yes, there is some leakage, but when I tried to fix that I got some funny results. Maybe my implementation wasn't right :? \n\n> -Did you try to sum mutiple previous responses ? Like a rolling window\n\nI didn't. I think that transformer can pick that up, right?",
    "1144378": "Simple yet effective ! Congratz on the strong finish !",
    "1144467": "Hi, congrats! Can you please elaborate on `max_quest` and `slide`?",
    "1144570": "This is related to the sampling strategy. `max_quest` is maximum number of questions processed at a time i.e. maximum input sequence length.\nNow, some users might have ` > max_quest` questions in their histories so you need to make multiple samples from one user. This is when `slide` comes into play. The simplest way to handle this is to set `slide = max_quest` and get samples from 0-300, 300-600, 600-900...\nBut this is suboptimal because questions close to the left hand side of the sample, like 300:350, or 600-650 can't see recent history. So I set `slide = 150` and make samples like this:\n\nsample1: 0-300 -> calculate loss on all positions\nsample2: 150-450 -> calculate loss on 300-450 positions\nsample3: 300-600 -> calculate loss on 450-600 positions\netc.\n\nHopefully this is clear enough.",
    "1144952": "Congrats and thank you for sharing!",
    "1145221": "Thanks a lot for the clear explanation. One more question related to `minute_embedding`, please. Is it `lag` and `prior_elapsed` converted to minutes, then multiplied by a shared learnable embedding?",
    "1145329": "thanks for sharing! a great job!",
    "1145843": "Yes, but embedding is not shared. Not that it makes too much of a difference;  I remember trying that and it giving me negligible worse result.",
    "1146286": "Simple and clever nice!! Regarding the inferenece did you strugle or had to optimize the code?\n\nEdit: any tips on this are welcome\n\nPs: I tried also last day to stack 5 weak models but didn't pass the inference stage and didn't have time to optimize it further (cv meta-model 0.805) - It was more as proof of concept and for learning rather than hitting the LB.. I should have tried with 3 though",
    "1146394": "Great, it was really fun to compete with you!",
    "1146528": "Ofcourse I struggled, out of first 13 submission, 4 were successful :) All three of my models preprocess the data in the same way, so I had to do it only once!\n\nI preprocessed the data locally into python dictionary with the following function, and uploaded it as a kaggle dataset. During inference, I was simply updating the inner dictionaries.\n\n```\ndef csv_to_dict(df):\n    \"\"\"maps the training data from csv to dictionary\"\"\"\n    # separate question events\n    questions_df = df[df[\"content_type_id\"] == False]\n        \n    # fill nans\n    questions_df[\"prior_question_elapsed_time\"] = questions_df[\"prior_question_elapsed_time\"].fillna(301000)\n        \n    # questions history container\n    questions_container = dict(tuple(questions_df.groupby([\"user_id\"])))\n    # df -> tensors\n    for user_id in questions_container.keys():\n        tmp_df = questions_container[user_id]\n        questions_container[user_id] = {\n            \"content_id\": torch.LongTensor(tmp_df[\"content_id\"].values),\n            \"timestamp\": torch.LongTensor(tmp_df[\"timestamp\"].values),\n            \"prior_question_elapsed_time\": torch.LongTensor(tmp_df[\"prior_question_elapsed_time\"].values),\n            \"answered_correctly\": torch.ShortTensor(tmp_df[\"answered_correctly\"].values),\n            }\n            \n    return questions_container\n\n```",
    "1146533": "Thanks, you were pushing me to limit, sorry I couldn't keep up :)  I was really tired last 2-3 weeks  (not that it's an excuse) and couldn't make almost no progress.",
    "1149597": "Congrats~~~ Your work is really impressive!\nI am a rookie in data science and I curiosity about how to generate 'question id embedding' features. Is it one-hot-encoding vector of question id?",
    "1149952": "No, it's a vector repsesentation: [Embedding](https://pytorch.org/docs/stable/generated/torch.nn.Embedding.html)",
    "1155694": "Wait. Did you say you only use 4 features? That's impressive!",
    "1155965": "Well who is learning, me or the model? :)"
  },
  "source": "meta"
}