{
  "id": 210715,
  "title": "126th place solution - LGBM with B-IRT features (LB 0.795)",
  "url": "/competitions/riiid-test-answer-prediction/discussion/210715",
  "author_name": "jwnz",
  "post_date": "2021-01-12T01:13:03.552000",
  "votes": 11,
  "comment_count": 0,
  "views": 0,
  "content": "<h4>Overview:</h4>\n<p>My solution is a single LightGBM model with a total of 48 features.</p>\n<h4>Best Features:</h4>\n<p>Unfortunately, I didn't keep track of exactly how much each feature affected my model's CV, but from empirical evidence I noticed the best features to be those based on:</p>\n<ul>\n<li>Elo rating system</li>\n<li>boolean determining whether or not a student has already seen a question</li>\n<li>previous timestamps</li>\n<li>Question difficulty based on the parameters of a Bayesian IRT model.</li>\n</ul>\n<p>A total list of the features is shown below (in no particular order):</p>\n<pre><code>analysis_columns = [\n    'answered_correctly',\n    'timestamp',\n    'prior_question_elapsed_time',\n    'prior_question_had_explanation',\n    'task_container_id',\n    'bundle_id',\n\n    'has_seen_question',\n    'user_correct_ratio',\n    'content_correct_ratio',\n    'correct_streak',\n    'incorrect_streak',\n    'user_total_question_count',\n    'part_correct_ratio',\n    'time_since_last_timestamp',\n    'time_since_last_lecture', \n    'tag_correct_ratio',\n    'previous_question_same_bundle',\n    'previous_correct_and_same_bundle',\n    'previous_lecture_type',\n    'previous_lecture_id',\n    'user_part_correct_ratio',\n    'user_part_avg_elapsed_time',\n    'user_part_timestamp_diff',\n    'user_tag_correct_ratio',\n    'user_tag_avg_elapsed_time',\n    'user_tag_timestamp_diff',\n    'question_history',\n    'lecture_count',\n    'questions_since_last_lec',\n    'hardest_question_correct_ratio',\n    'easiest_question_correct_ratio',\n    'smallest_pqet',\n    'largest_pqet',\n    'user_rating',\n    'content_rating',\n    'highest_elo',\n    'lowest_elo',\n    'elo_user_content',\n    'elo_content_user',\n    'tag',\n    'irt_difficulty',\n    'irt_discrimination',\n    'timestamp_n1',\n    'timestamp_n2', \n    'timestamp_n3',\n    'tag_difficulty',\n    'tag_discrimination',\n    'pqhe_true_correct_ratio',\n    'pqhe_false_correct_ratio',\n]\n</code></pre>\n<p>Feature importance graph (Note: the names somewhat vary from above):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3785972%2F3078d4cb7724c798c7091f7dbb48f66e%2F.png?generation=1610411000057392&amp;alt=media\" alt=\"\"></p>\n<h4>Things I tried that didn't work:</h4>\n<p><strong>Factorization Machine:</strong><br>\nAs others have stated, the goal of this task can be, to some extent, compared to that of a recommendation system where the recommendation score can be interpreted as a student's ability to answer some question. Although not without criticism, Riiid! has also used or referenced Matrix Factorization in their own work (see <a href=\"https://educationaldatamining.org/files/conferences/EDM2020/papers/paper_91.pdf\" target=\"_blank\">[1]</a>, <a href=\"https://arxiv.org/pdf/2005.05021.pdf\" target=\"_blank\">[2]</a>, <a href=\"https://arxiv.org/pdf/2012.05031.pdf\" target=\"_blank\">[3]</a>).</p>\n<p>The <a href=\"https://ieeexplore.ieee.org/document/5694074\" target=\"_blank\">Factorization Machine</a> algorithm is quite simple as is as follows:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3785972%2F137e8b1e4d9cb771c1c526f665801e65%2Fsagemaker-factorization-1.png?generation=1610412377437882&amp;alt=media\" alt=\"\"></p>\n<p>Unfortunately, I was unable to get a successful submission using existing Factorization Machine tools, as well as my own implementation. Instead, I tried incorporating the latent factors from the <em>V</em> interaction matrix (shown in the formula above), but they didn't have any noticeable positive effect on the outcome of the model.</p>\n<p><strong>Bayesian Knowledge Tracing (BKT):</strong><br>\nBayesian Knowledge Tracing is the precursor to Deep Knowledge Tracing. The particular model I was interested in was the individualized variation that include contextualization of guessing and slipping estimates. My goal was to train a model offline and use the <em>guess</em> and <em>slip</em> parameters in the LGBM model. Unfortunately, these also showed very little positive effect on the model's overall CV. This may be due to the poor performance of the trained BKT model as well.</p>\n<p><strong>Neural Halflife Regression:</strong><br>\nUsing Tensorflow, I also intended to implement forgetting of a question's tag into my model based on Duolingo's <a href=\"https://research.duolingo.com/papers/settles.acl16.pdf\" target=\"_blank\">Halflife regression</a> model, inspired by <a href=\"https://link.springer.com/chapter/10.1007/978-3-030-52240-7_65\" target=\"_blank\">this</a> paper.</p>\n<p>The formula used is also quite simple and is as follows:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3785972%2Fb8e08c0ede4d0a3244dad4a882ea8cba%2F2021-01-12%20100012.png?generation=1610413465690518&amp;alt=media\" alt=\"\"><br>\n, where <em>h</em> and <em>C</em>, are encoders, representing a learner's strength and a question's complexity respectively.</p>\n<p>However, in the short amount of time I had left, I was unable to ensemble the model with the LGBM model to test whether or not it may have had a positive effect on the outcome of the model.</p>\n<h4>Conclusion</h4>\n<p>Overall this was a great experience, I learned a lot and had a lot of fun. I hope to get more comfortable with TensorFlow and PyTorch so I can have a more competitive edge in future competitions.</p>\n<p>All of the code I used for this project I have uploaded to the following github repository: <a href=\"https://github.com/jwnz/Riiid-AIEd-Challenge-2020\" target=\"_blank\">https://github.com/jwnz/Riiid-AIEd-Challenge-2020</a></p>",
  "messages": [
    {
      "id": 1149579,
      "postDate": "2021-01-12T01:13:03.553Z",
      "content": "<h4>Overview:</h4>\n<p>My solution is a single LightGBM model with a total of 48 features.</p>\n<h4>Best Features:</h4>\n<p>Unfortunately, I didn't keep track of exactly how much each feature affected my model's CV, but from empirical evidence I noticed the best features to be those based on:</p>\n<ul>\n<li>Elo rating system</li>\n<li>boolean determining whether or not a student has already seen a question</li>\n<li>previous timestamps</li>\n<li>Question difficulty based on the parameters of a Bayesian IRT model.</li>\n</ul>\n<p>A total list of the features is shown below (in no particular order):</p>\n<pre><code>analysis_columns = [\n    'answered_correctly',\n    'timestamp',\n    'prior_question_elapsed_time',\n    'prior_question_had_explanation',\n    'task_container_id',\n    'bundle_id',\n\n    'has_seen_question',\n    'user_correct_ratio',\n    'content_correct_ratio',\n    'correct_streak',\n    'incorrect_streak',\n    'user_total_question_count',\n    'part_correct_ratio',\n    'time_since_last_timestamp',\n    'time_since_last_lecture', \n    'tag_correct_ratio',\n    'previous_question_same_bundle',\n    'previous_correct_and_same_bundle',\n    'previous_lecture_type',\n    'previous_lecture_id',\n    'user_part_correct_ratio',\n    'user_part_avg_elapsed_time',\n    'user_part_timestamp_diff',\n    'user_tag_correct_ratio',\n    'user_tag_avg_elapsed_time',\n    'user_tag_timestamp_diff',\n    'question_history',\n    'lecture_count',\n    'questions_since_last_lec',\n    'hardest_question_correct_ratio',\n    'easiest_question_correct_ratio',\n    'smallest_pqet',\n    'largest_pqet',\n    'user_rating',\n    'content_rating',\n    'highest_elo',\n    'lowest_elo',\n    'elo_user_content',\n    'elo_content_user',\n    'tag',\n    'irt_difficulty',\n    'irt_discrimination',\n    'timestamp_n1',\n    'timestamp_n2', \n    'timestamp_n3',\n    'tag_difficulty',\n    'tag_discrimination',\n    'pqhe_true_correct_ratio',\n    'pqhe_false_correct_ratio',\n]\n</code></pre>\n<p>Feature importance graph (Note: the names somewhat vary from above):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3785972%2F3078d4cb7724c798c7091f7dbb48f66e%2F.png?generation=1610411000057392&amp;alt=media\" alt=\"\"></p>\n<h4>Things I tried that didn't work:</h4>\n<p><strong>Factorization Machine:</strong><br>\nAs others have stated, the goal of this task can be, to some extent, compared to that of a recommendation system where the recommendation score can be interpreted as a student's ability to answer some question. Although not without criticism, Riiid! has also used or referenced Matrix Factorization in their own work (see <a href=\"https://educationaldatamining.org/files/conferences/EDM2020/papers/paper_91.pdf\" target=\"_blank\">[1]</a>, <a href=\"https://arxiv.org/pdf/2005.05021.pdf\" target=\"_blank\">[2]</a>, <a href=\"https://arxiv.org/pdf/2012.05031.pdf\" target=\"_blank\">[3]</a>).</p>\n<p>The <a href=\"https://ieeexplore.ieee.org/document/5694074\" target=\"_blank\">Factorization Machine</a> algorithm is quite simple as is as follows:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3785972%2F137e8b1e4d9cb771c1c526f665801e65%2Fsagemaker-factorization-1.png?generation=1610412377437882&amp;alt=media\" alt=\"\"></p>\n<p>Unfortunately, I was unable to get a successful submission using existing Factorization Machine tools, as well as my own implementation. Instead, I tried incorporating the latent factors from the <em>V</em> interaction matrix (shown in the formula above), but they didn't have any noticeable positive effect on the outcome of the model.</p>\n<p><strong>Bayesian Knowledge Tracing (BKT):</strong><br>\nBayesian Knowledge Tracing is the precursor to Deep Knowledge Tracing. The particular model I was interested in was the individualized variation that include contextualization of guessing and slipping estimates. My goal was to train a model offline and use the <em>guess</em> and <em>slip</em> parameters in the LGBM model. Unfortunately, these also showed very little positive effect on the model's overall CV. This may be due to the poor performance of the trained BKT model as well.</p>\n<p><strong>Neural Halflife Regression:</strong><br>\nUsing Tensorflow, I also intended to implement forgetting of a question's tag into my model based on Duolingo's <a href=\"https://research.duolingo.com/papers/settles.acl16.pdf\" target=\"_blank\">Halflife regression</a> model, inspired by <a href=\"https://link.springer.com/chapter/10.1007/978-3-030-52240-7_65\" target=\"_blank\">this</a> paper.</p>\n<p>The formula used is also quite simple and is as follows:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3785972%2Fb8e08c0ede4d0a3244dad4a882ea8cba%2F2021-01-12%20100012.png?generation=1610413465690518&amp;alt=media\" alt=\"\"><br>\n, where <em>h</em> and <em>C</em>, are encoders, representing a learner's strength and a question's complexity respectively.</p>\n<p>However, in the short amount of time I had left, I was unable to ensemble the model with the LGBM model to test whether or not it may have had a positive effect on the outcome of the model.</p>\n<h4>Conclusion</h4>\n<p>Overall this was a great experience, I learned a lot and had a lot of fun. I hope to get more comfortable with TensorFlow and PyTorch so I can have a more competitive edge in future competitions.</p>\n<p>All of the code I used for this project I have uploaded to the following github repository: <a href=\"https://github.com/jwnz/Riiid-AIEd-Challenge-2020\" target=\"_blank\">https://github.com/jwnz/Riiid-AIEd-Challenge-2020</a></p>",
      "rawMarkdown": "#### Overview:\n\nMy solution is a single LightGBM model with a total of 48 features.\n\n\n#### Best Features:\n\nUnfortunately, I didn't keep track of exactly how much each feature affected my model's CV, but from empirical evidence I noticed the best features to be those based on:\n- Elo rating system\n- boolean determining whether or not a student has already seen a question\n- previous timestamps\n- Question difficulty based on the parameters of a Bayesian IRT model.\n\nA total list of the features is shown below (in no particular order):\n\n```\nanalysis_columns = [\n    'answered_correctly',\n    'timestamp',\n    'prior_question_elapsed_time',\n    'prior_question_had_explanation',\n    'task_container_id',\n    'bundle_id',\n     \n    'has_seen_question',\n    'user_correct_ratio',\n    'content_correct_ratio',\n    'correct_streak',\n    'incorrect_streak',\n    'user_total_question_count',\n    'part_correct_ratio',\n    'time_since_last_timestamp',\n    'time_since_last_lecture', \n    'tag_correct_ratio',\n    'previous_question_same_bundle',\n    'previous_correct_and_same_bundle',\n    'previous_lecture_type',\n    'previous_lecture_id',\n    'user_part_correct_ratio',\n    'user_part_avg_elapsed_time',\n    'user_part_timestamp_diff',\n    'user_tag_correct_ratio',\n    'user_tag_avg_elapsed_time',\n    'user_tag_timestamp_diff',\n    'question_history',\n    'lecture_count',\n    'questions_since_last_lec',\n    'hardest_question_correct_ratio',\n    'easiest_question_correct_ratio',\n    'smallest_pqet',\n    'largest_pqet',\n    'user_rating',\n    'content_rating',\n    'highest_elo',\n    'lowest_elo',\n    'elo_user_content',\n    'elo_content_user',\n    'tag',\n    'irt_difficulty',\n    'irt_discrimination',\n    'timestamp_n1',\n    'timestamp_n2', \n    'timestamp_n3',\n    'tag_difficulty',\n    'tag_discrimination',\n    'pqhe_true_correct_ratio',\n    'pqhe_false_correct_ratio',\n]\n```\n\nFeature importance graph (Note: the names somewhat vary from above):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3785972%2F3078d4cb7724c798c7091f7dbb48f66e%2F.png?generation=1610411000057392&alt=media)\n\n\n#### Things I tried that didn't work:\n\n**Factorization Machine:**\nAs others have stated, the goal of this task can be, to some extent, compared to that of a recommendation system where the recommendation score can be interpreted as a student's ability to answer some question. Although not without criticism, Riiid! has also used or referenced Matrix Factorization in their own work (see [[1]](https://educationaldatamining.org/files/conferences/EDM2020/papers/paper_91.pdf), [[2]](https://arxiv.org/pdf/2005.05021.pdf), [[3]](https://arxiv.org/pdf/2012.05031.pdf)).\n\nThe [Factorization Machine](https://ieeexplore.ieee.org/document/5694074) algorithm is quite simple as is as follows:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3785972%2F137e8b1e4d9cb771c1c526f665801e65%2Fsagemaker-factorization-1.png?generation=1610412377437882&alt=media)\n\nUnfortunately, I was unable to get a successful submission using existing Factorization Machine tools, as well as my own implementation. Instead, I tried incorporating the latent factors from the *V* interaction matrix (shown in the formula above), but they didn't have any noticeable positive effect on the outcome of the model.\n\n**Bayesian Knowledge Tracing (BKT):**\nBayesian Knowledge Tracing is the precursor to Deep Knowledge Tracing. The particular model I was interested in was the individualized variation that include contextualization of guessing and slipping estimates. My goal was to train a model offline and use the *guess* and *slip* parameters in the LGBM model. Unfortunately, these also showed very little positive effect on the model's overall CV. This may be due to the poor performance of the trained BKT model as well.\n\n**Neural Halflife Regression:**\nUsing Tensorflow, I also intended to implement forgetting of a question's tag into my model based on Duolingo's [Halflife regression](https://research.duolingo.com/papers/settles.acl16.pdf) model, inspired by [this](https://link.springer.com/chapter/10.1007/978-3-030-52240-7_65) paper.\n\nThe formula used is also quite simple and is as follows:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3785972%2Fb8e08c0ede4d0a3244dad4a882ea8cba%2F2021-01-12%20100012.png?generation=1610413465690518&alt=media)\n, where *h* and *C*, are encoders, representing a learner's strength and a question's complexity respectively.\n\nHowever, in the short amount of time I had left, I was unable to ensemble the model with the LGBM model to test whether or not it may have had a positive effect on the outcome of the model.\n\n\n#### Conclusion\n\nOverall this was a great experience, I learned a lot and had a lot of fun. I hope to get more comfortable with TensorFlow and PyTorch so I can have a more competitive edge in future competitions.\n\nAll of the code I used for this project I have uploaded to the following github repository: https://github.com/jwnz/Riiid-AIEd-Challenge-2020",
      "votes": 11
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1149579": "#### Overview:\n\nMy solution is a single LightGBM model with a total of 48 features.\n\n\n#### Best Features:\n\nUnfortunately, I didn't keep track of exactly how much each feature affected my model's CV, but from empirical evidence I noticed the best features to be those based on:\n- Elo rating system\n- boolean determining whether or not a student has already seen a question\n- previous timestamps\n- Question difficulty based on the parameters of a Bayesian IRT model.\n\nA total list of the features is shown below (in no particular order):\n\n```\nanalysis_columns = [\n    'answered_correctly',\n    'timestamp',\n    'prior_question_elapsed_time',\n    'prior_question_had_explanation',\n    'task_container_id',\n    'bundle_id',\n     \n    'has_seen_question',\n    'user_correct_ratio',\n    'content_correct_ratio',\n    'correct_streak',\n    'incorrect_streak',\n    'user_total_question_count',\n    'part_correct_ratio',\n    'time_since_last_timestamp',\n    'time_since_last_lecture', \n    'tag_correct_ratio',\n    'previous_question_same_bundle',\n    'previous_correct_and_same_bundle',\n    'previous_lecture_type',\n    'previous_lecture_id',\n    'user_part_correct_ratio',\n    'user_part_avg_elapsed_time',\n    'user_part_timestamp_diff',\n    'user_tag_correct_ratio',\n    'user_tag_avg_elapsed_time',\n    'user_tag_timestamp_diff',\n    'question_history',\n    'lecture_count',\n    'questions_since_last_lec',\n    'hardest_question_correct_ratio',\n    'easiest_question_correct_ratio',\n    'smallest_pqet',\n    'largest_pqet',\n    'user_rating',\n    'content_rating',\n    'highest_elo',\n    'lowest_elo',\n    'elo_user_content',\n    'elo_content_user',\n    'tag',\n    'irt_difficulty',\n    'irt_discrimination',\n    'timestamp_n1',\n    'timestamp_n2', \n    'timestamp_n3',\n    'tag_difficulty',\n    'tag_discrimination',\n    'pqhe_true_correct_ratio',\n    'pqhe_false_correct_ratio',\n]\n```\n\nFeature importance graph (Note: the names somewhat vary from above):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3785972%2F3078d4cb7724c798c7091f7dbb48f66e%2F.png?generation=1610411000057392&alt=media)\n\n\n#### Things I tried that didn't work:\n\n**Factorization Machine:**\nAs others have stated, the goal of this task can be, to some extent, compared to that of a recommendation system where the recommendation score can be interpreted as a student's ability to answer some question. Although not without criticism, Riiid! has also used or referenced Matrix Factorization in their own work (see [[1]](https://educationaldatamining.org/files/conferences/EDM2020/papers/paper_91.pdf), [[2]](https://arxiv.org/pdf/2005.05021.pdf), [[3]](https://arxiv.org/pdf/2012.05031.pdf)).\n\nThe [Factorization Machine](https://ieeexplore.ieee.org/document/5694074) algorithm is quite simple as is as follows:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3785972%2F137e8b1e4d9cb771c1c526f665801e65%2Fsagemaker-factorization-1.png?generation=1610412377437882&alt=media)\n\nUnfortunately, I was unable to get a successful submission using existing Factorization Machine tools, as well as my own implementation. Instead, I tried incorporating the latent factors from the *V* interaction matrix (shown in the formula above), but they didn't have any noticeable positive effect on the outcome of the model.\n\n**Bayesian Knowledge Tracing (BKT):**\nBayesian Knowledge Tracing is the precursor to Deep Knowledge Tracing. The particular model I was interested in was the individualized variation that include contextualization of guessing and slipping estimates. My goal was to train a model offline and use the *guess* and *slip* parameters in the LGBM model. Unfortunately, these also showed very little positive effect on the model's overall CV. This may be due to the poor performance of the trained BKT model as well.\n\n**Neural Halflife Regression:**\nUsing Tensorflow, I also intended to implement forgetting of a question's tag into my model based on Duolingo's [Halflife regression](https://research.duolingo.com/papers/settles.acl16.pdf) model, inspired by [this](https://link.springer.com/chapter/10.1007/978-3-030-52240-7_65) paper.\n\nThe formula used is also quite simple and is as follows:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3785972%2Fb8e08c0ede4d0a3244dad4a882ea8cba%2F2021-01-12%20100012.png?generation=1610413465690518&alt=media)\n, where *h* and *C*, are encoders, representing a learner's strength and a question's complexity respectively.\n\nHowever, in the short amount of time I had left, I was unable to ensemble the model with the LGBM model to test whether or not it may have had a positive effect on the outcome of the model.\n\n\n#### Conclusion\n\nOverall this was a great experience, I learned a lot and had a lot of fun. I hope to get more comfortable with TensorFlow and PyTorch so I can have a more competitive edge in future competitions.\n\nAll of the code I used for this project I have uploaded to the following github repository: https://github.com/jwnz/Riiid-AIEd-Challenge-2020"
  }
}