{
  "id": 212090,
  "title": "149th solution write up (LGBM + Transformer ensemble)",
  "url": "/competitions/riiid-test-answer-prediction/writeups/1-149th-solution-write-up-lgbm-transformer-ensembl",
  "author_name": "",
  "post_date": "2021-01-17T14:40:00.603Z",
  "votes": 10,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Also available at <a href=\"https://github.com/yoonseok312/riiid-answer-correctness-prediction\" target=\"_blank\">https://github.com/yoonseok312/riiid-answer-correctness-prediction</a> with codes.</p>\n<p>This is a solution write up for our model, which is an ensemble between a single Light Gradient Boosted Machine model, and a single Encoder-Decoder based Transformer model. It was our first time competing in a Kaggle competition and none of us had previous AI/ML/Stats experience so we learned a lot throughout this competition. Should you have any questions, feel free to contact me at yoonseok@berkeley.edu.</p>\n<h1>LightGBM</h1>\n<h2>Features</h2>\n<p>ts_delta: the gap between timestamp of current content with previous content of the same user.<br>\ntask_container_id<br>\nprior_question_elapsed_time: in second, rounded<br>\nprior_question_had_explanation<br>\npart<br>\nnum_tag: number of tags in that question<br>\nu_chance: the average correctness of the user until current time<br>\nu_attempts: number of content the user have done<br>\nu_attempt_c: number of times the user interacted with the specific content in the past (only counting from &gt;1 interactions due to memory)<br>\nc_chance: the average correctness of the question until current time<br>\nc_attempts: number of encounter of that question (all user)<br>\nu_part_chance: the average correctness of the user doing the same part as the question<br>\nu_part_attempts: number of question of the same part the user have done<br>\nu_skill_chance: the average correctness of the user doing the same skill as the question (part &lt; 5: listening, part &gt;= 5: reading)<br>\nu_skill_attempts: number of question of the same skill the user have done<br>\nt_chance: the average correctness of the user of questions with specific tag until current time<br>\nt_attempts: user's number of encounter of that tag<br>\ntotal_explained: number of times explanation was provided to the user until current time (all contents)<br>\n10_recent_correctness: user correctness of the most recent questions (up to 10)<br>\n10_recent_mean_gap: mean ts_delta of the most recent questions (up to 10)<br>\nbundle_elapsed: the mean elapsed time of the bundle, up until the abs time.<br>\nmean_elapsed: the mean elapsed time of the user until now.<br>\nprev_t1: tag of the last question<br>\nprev_cor: correctness of the last question<br>\ntrueskill_possibility: possibility of the user 'beating' the question (getting the question correct) based on trueskill<br>\nmu: mu value (mean of trueskill ratings) of user<br>\nsigma: sigma value (standard deviation of trueskill ratings) of user<br>\nColumns with NaN value was filled with -1.</p>\n<h2>Cross Validation and train strategy</h2>\n<ol>\n<li>Define an absolute time for the whole database (abs_time = user_id//50 + timestamp//1000)</li>\n<li>Sort by abs_time. All features mentioned above were engineered such that we will not take data from the future (higher abs_time) into account.</li>\n<li>Drop first 25% of the data. This data contains some noise, i.e. when nobody studied plenty of questions yet.</li>\n<li>Take last 25% of data as validation set.</li>\n<li>Train the model with remaining 50% of the data.</li>\n</ol>\n<h2>Single Model AUC</h2>\n<p>We used less than 30 features, but considering that most of the single LGBM models above 0.79 AUC used 40+ features, we did a decent work on focusing on imoportant features.<br>\nLB score: AUC 0.789<br>\nNumber of epochs: 6650 (around 15 hours of training in total)</p>\n<h1>Transformer</h1>\n<h2>Encoder</h2>\n<p>Added below layers with positional encoding.</p>\n<p>1) Excercise Related<br>\nmin_delta: minute difference from between this question and the previous. Cap at 1443 (1 day)<br>\nday_delta: day difference from between this question and the previous. Cap at 30<br>\nmonth_delta: month difference from between this question and the previous. Cap at 9<br>\ntid: task container id<br>\nis_with: if the question is presented with another question. Usually have the same task container<br>\nc_part: part, one hot encode and denote skill (listening, reading, part1,2,…)<br>\ntag1…6: tags of question (t1 to t6 are the tag of one question, t1 being the most important tag.) Above embeddings or Dense layer concatenated.</p>\n<p>2) Content id (cid)<br>\nDense layer</p>\n<h2>Decoder</h2>\n<p>Added below layers with positional encoding.</p>\n<p>1) Response Related<br>\nprev_answered_correct: correctness of previous answer.<br>\nprior_elapsed: prior elapsed tiem<br>\nprior_explained: prior has explanation Above embeddings or Dense layer concatenated.</p>\n<p>2) Answered Correctly<br>\nConcatenate Lecture related Embeddings/Dense<br>\nnum_lect: number of lecture the user have seen<br>\nlec_type: 1 hot encode of most recent lecture type, (llecty1, 2…)<br>\nlec_h_past: time since most recent lecture Above embeddings or Dense layer concatenated.</p>\n<h2>Parameters</h2>\n<p>WINDOW_SIZE: 100<br>\nEMBED_DIM: 256<br>\nNUM_HEADS: 16</p>\n<h2>Cross Validation and train strategy</h2>\n<p>Use first 80% of data as train set and last 20% as validation set.</p>\n<h2>Single Model AUC</h2>\n<p>AUC 0.786<br>\nSAINT model has plenty of room for improvement, but as 1 epoch took more than 10 hours to train we decided to focus on improving LGBM.</p>\n<h1>Inference</h1>\n<h2>1. Ensembling two models</h2>\n<p>Ensembled a single LGBM model and a Transformer model in 0.55 (LGBM) / 0.45 ratio.<br>\nAUC: 0.793</p>\n<h2>2. Ensembling three models</h2>\n<p>Ensembled two LGBM and a Transformer model. 2nd LGBM was same as the first LGBM but except features related to Trueskill. When 2 models out of three models predicted that the user is likely to answer correctly, we used the max value among the 3 predictions. When 2 models out of three models predicted that the user is likely to answer wronly, we used the min value among the 3 predictions. For remaining cases, we mixed three models in 0.4 (Transformer) / 0.45 (First LGBM) / 0.15 (Second LGBM) ratio.<br>\nAUC: 0.793 (slightly higher than the 1st Inference)</p>\n<h1>Training Environment</h1>\n<p>Our biggest mistake was thinking that all the feature engineering, training, and inferencing process must be done in the Kaggle environment. We were only using Kaggle environment until 2 weeks before the competiton ended, and from then we started to use Google Colab with GPU and 25GB of RAM. Still, there were several times when Colab took GPU from us and didn't give it for several hours as we were constanly using their GPU.</p>\n<p>Thanks to my teammates <a href=\"https://www.kaggle.com/shhrkre\" target=\"_blank\">@shhrkre</a>, <a href=\"https://www.kaggle.com/ysgong\" target=\"_blank\">@ysgong</a>, <a href=\"https://www.kaggle.com/sanmaruum\" target=\"_blank\">@sanmaruum</a>, <a href=\"https://www.kaggle.com/kuraji\" target=\"_blank\">@kuraji</a>.</p>",
  "messages": [
    {
      "id": "1156923",
      "postDate": "01/17/2021 14:21:34",
      "content": "<p>Also available at <a href=\"https://github.com/yoonseok312/riiid-answer-correctness-prediction\" target=\"_blank\">https://github.com/yoonseok312/riiid-answer-correctness-prediction</a> with codes.</p>\n<p>This is a solution write up for our model, which is an ensemble between a single Light Gradient Boosted Machine model, and a single Encoder-Decoder based Transformer model. It was our first time competing in a Kaggle competition and none of us had previous AI/ML/Stats experience so we learned a lot throughout this competition. Should you have any questions, feel free to contact me at yoonseok@berkeley.edu.</p>\n<h1>LightGBM</h1>\n<h2>Features</h2>\n<p>ts_delta: the gap between timestamp of current content with previous content of the same user.<br>\ntask_container_id<br>\nprior_question_elapsed_time: in second, rounded<br>\nprior_question_had_explanation<br>\npart<br>\nnum_tag: number of tags in that question<br>\nu_chance: the average correctness of the user until current time<br>\nu_attempts: number of content the user have done<br>\nu_attempt_c: number of times the user interacted with the specific content in the past (only counting from &gt;1 interactions due to memory)<br>\nc_chance: the average correctness of the question until current time<br>\nc_attempts: number of encounter of that question (all user)<br>\nu_part_chance: the average correctness of the user doing the same part as the question<br>\nu_part_attempts: number of question of the same part the user have done<br>\nu_skill_chance: the average correctness of the user doing the same skill as the question (part &lt; 5: listening, part &gt;= 5: reading)<br>\nu_skill_attempts: number of question of the same skill the user have done<br>\nt_chance: the average correctness of the user of questions with specific tag until current time<br>\nt_attempts: user's number of encounter of that tag<br>\ntotal_explained: number of times explanation was provided to the user until current time (all contents)<br>\n10_recent_correctness: user correctness of the most recent questions (up to 10)<br>\n10_recent_mean_gap: mean ts_delta of the most recent questions (up to 10)<br>\nbundle_elapsed: the mean elapsed time of the bundle, up until the abs time.<br>\nmean_elapsed: the mean elapsed time of the user until now.<br>\nprev_t1: tag of the last question<br>\nprev_cor: correctness of the last question<br>\ntrueskill_possibility: possibility of the user 'beating' the question (getting the question correct) based on trueskill<br>\nmu: mu value (mean of trueskill ratings) of user<br>\nsigma: sigma value (standard deviation of trueskill ratings) of user<br>\nColumns with NaN value was filled with -1.</p>\n<h2>Cross Validation and train strategy</h2>\n<ol>\n<li>Define an absolute time for the whole database (abs_time = user_id//50 + timestamp//1000)</li>\n<li>Sort by abs_time. All features mentioned above were engineered such that we will not take data from the future (higher abs_time) into account.</li>\n<li>Drop first 25% of the data. This data contains some noise, i.e. when nobody studied plenty of questions yet.</li>\n<li>Take last 25% of data as validation set.</li>\n<li>Train the model with remaining 50% of the data.</li>\n</ol>\n<h2>Single Model AUC</h2>\n<p>We used less than 30 features, but considering that most of the single LGBM models above 0.79 AUC used 40+ features, we did a decent work on focusing on imoportant features.<br>\nLB score: AUC 0.789<br>\nNumber of epochs: 6650 (around 15 hours of training in total)</p>\n<h1>Transformer</h1>\n<h2>Encoder</h2>\n<p>Added below layers with positional encoding.</p>\n<p>1) Excercise Related<br>\nmin_delta: minute difference from between this question and the previous. Cap at 1443 (1 day)<br>\nday_delta: day difference from between this question and the previous. Cap at 30<br>\nmonth_delta: month difference from between this question and the previous. Cap at 9<br>\ntid: task container id<br>\nis_with: if the question is presented with another question. Usually have the same task container<br>\nc_part: part, one hot encode and denote skill (listening, reading, part1,2,…)<br>\ntag1…6: tags of question (t1 to t6 are the tag of one question, t1 being the most important tag.) Above embeddings or Dense layer concatenated.</p>\n<p>2) Content id (cid)<br>\nDense layer</p>\n<h2>Decoder</h2>\n<p>Added below layers with positional encoding.</p>\n<p>1) Response Related<br>\nprev_answered_correct: correctness of previous answer.<br>\nprior_elapsed: prior elapsed tiem<br>\nprior_explained: prior has explanation Above embeddings or Dense layer concatenated.</p>\n<p>2) Answered Correctly<br>\nConcatenate Lecture related Embeddings/Dense<br>\nnum_lect: number of lecture the user have seen<br>\nlec_type: 1 hot encode of most recent lecture type, (llecty1, 2…)<br>\nlec_h_past: time since most recent lecture Above embeddings or Dense layer concatenated.</p>\n<h2>Parameters</h2>\n<p>WINDOW_SIZE: 100<br>\nEMBED_DIM: 256<br>\nNUM_HEADS: 16</p>\n<h2>Cross Validation and train strategy</h2>\n<p>Use first 80% of data as train set and last 20% as validation set.</p>\n<h2>Single Model AUC</h2>\n<p>AUC 0.786<br>\nSAINT model has plenty of room for improvement, but as 1 epoch took more than 10 hours to train we decided to focus on improving LGBM.</p>\n<h1>Inference</h1>\n<h2>1. Ensembling two models</h2>\n<p>Ensembled a single LGBM model and a Transformer model in 0.55 (LGBM) / 0.45 ratio.<br>\nAUC: 0.793</p>\n<h2>2. Ensembling three models</h2>\n<p>Ensembled two LGBM and a Transformer model. 2nd LGBM was same as the first LGBM but except features related to Trueskill. When 2 models out of three models predicted that the user is likely to answer correctly, we used the max value among the 3 predictions. When 2 models out of three models predicted that the user is likely to answer wronly, we used the min value among the 3 predictions. For remaining cases, we mixed three models in 0.4 (Transformer) / 0.45 (First LGBM) / 0.15 (Second LGBM) ratio.<br>\nAUC: 0.793 (slightly higher than the 1st Inference)</p>\n<h1>Training Environment</h1>\n<p>Our biggest mistake was thinking that all the feature engineering, training, and inferencing process must be done in the Kaggle environment. We were only using Kaggle environment until 2 weeks before the competiton ended, and from then we started to use Google Colab with GPU and 25GB of RAM. Still, there were several times when Colab took GPU from us and didn't give it for several hours as we were constanly using their GPU.</p>\n<p>Thanks to my teammates <a href=\"https://www.kaggle.com/shhrkre\" target=\"_blank\">@shhrkre</a>, <a href=\"https://www.kaggle.com/ysgong\" target=\"_blank\">@ysgong</a>, <a href=\"https://www.kaggle.com/sanmaruum\" target=\"_blank\">@sanmaruum</a>, <a href=\"https://www.kaggle.com/kuraji\" target=\"_blank\">@kuraji</a>.</p>",
      "rawMarkdown": "Also available at https://github.com/yoonseok312/riiid-answer-correctness-prediction with codes.\n\nThis is a solution write up for our model, which is an ensemble between a single Light Gradient Boosted Machine model, and a single Encoder-Decoder based Transformer model. It was our first time competing in a Kaggle competition and none of us had previous AI/ML/Stats experience so we learned a lot throughout this competition. Should you have any questions, feel free to contact me at yoonseok@berkeley.edu.\n\n# LightGBM\n## Features\nts_delta: the gap between timestamp of current content with previous content of the same user.\ntask_container_id\nprior_question_elapsed_time: in second, rounded\nprior_question_had_explanation\npart\nnum_tag: number of tags in that question\nu_chance: the average correctness of the user until current time\nu_attempts: number of content the user have done\nu_attempt_c: number of times the user interacted with the specific content in the past (only counting from >1 interactions due to memory)\nc_chance: the average correctness of the question until current time\nc_attempts: number of encounter of that question (all user)\nu_part_chance: the average correctness of the user doing the same part as the question\nu_part_attempts: number of question of the same part the user have done\nu_skill_chance: the average correctness of the user doing the same skill as the question (part < 5: listening, part >= 5: reading)\nu_skill_attempts: number of question of the same skill the user have done\nt_chance: the average correctness of the user of questions with specific tag until current time\nt_attempts: user's number of encounter of that tag\ntotal_explained: number of times explanation was provided to the user until current time (all contents)\n10_recent_correctness: user correctness of the most recent questions (up to 10)\n10_recent_mean_gap: mean ts_delta of the most recent questions (up to 10)\nbundle_elapsed: the mean elapsed time of the bundle, up until the abs time.\nmean_elapsed: the mean elapsed time of the user until now.\nprev_t1: tag of the last question\nprev_cor: correctness of the last question\ntrueskill_possibility: possibility of the user 'beating' the question (getting the question correct) based on trueskill\nmu: mu value (mean of trueskill ratings) of user\nsigma: sigma value (standard deviation of trueskill ratings) of user\nColumns with NaN value was filled with -1.\n\n## Cross Validation and train strategy\n1. Define an absolute time for the whole database (abs_time = user_id//50 + timestamp//1000)\n2. Sort by abs_time. All features mentioned above were engineered such that we will not take data from the future (higher abs_time) into account.\n3. Drop first 25% of the data. This data contains some noise, i.e. when nobody studied plenty of questions yet.\n4. Take last 25% of data as validation set.\n5. Train the model with remaining 50% of the data.\n\n## Single Model AUC\nWe used less than 30 features, but considering that most of the single LGBM models above 0.79 AUC used 40+ features, we did a decent work on focusing on imoportant features.\nLB score: AUC 0.789\nNumber of epochs: 6650 (around 15 hours of training in total)\n\n# Transformer\n## Encoder\nAdded below layers with positional encoding.\n\n1) Excercise Related\nmin_delta: minute difference from between this question and the previous. Cap at 1443 (1 day)\nday_delta: day difference from between this question and the previous. Cap at 30\nmonth_delta: month difference from between this question and the previous. Cap at 9\ntid: task container id\nis_with: if the question is presented with another question. Usually have the same task container\nc_part: part, one hot encode and denote skill (listening, reading, part1,2,...)\ntag1...6: tags of question (t1 to t6 are the tag of one question, t1 being the most important tag.) Above embeddings or Dense layer concatenated.\n\n2) Content id (cid)\nDense layer\n\n## Decoder\nAdded below layers with positional encoding.\n\n1) Response Related\nprev_answered_correct: correctness of previous answer.\nprior_elapsed: prior elapsed tiem\nprior_explained: prior has explanation Above embeddings or Dense layer concatenated.\n\n2) Answered Correctly\nConcatenate Lecture related Embeddings/Dense\nnum_lect: number of lecture the user have seen\nlec_type: 1 hot encode of most recent lecture type, (llecty1, 2...)\nlec_h_past: time since most recent lecture Above embeddings or Dense layer concatenated.\n\n## Parameters\nWINDOW_SIZE: 100\nEMBED_DIM: 256\nNUM_HEADS: 16\n\n## Cross Validation and train strategy\nUse first 80% of data as train set and last 20% as validation set.\n\n## Single Model AUC\nAUC 0.786\nSAINT model has plenty of room for improvement, but as 1 epoch took more than 10 hours to train we decided to focus on improving LGBM.\n\n# Inference\n## 1. Ensembling two models\nEnsembled a single LGBM model and a Transformer model in 0.55 (LGBM) / 0.45 ratio.\nAUC: 0.793\n\n## 2. Ensembling three models\nEnsembled two LGBM and a Transformer model. 2nd LGBM was same as the first LGBM but except features related to Trueskill. When 2 models out of three models predicted that the user is likely to answer correctly, we used the max value among the 3 predictions. When 2 models out of three models predicted that the user is likely to answer wronly, we used the min value among the 3 predictions. For remaining cases, we mixed three models in 0.4 (Transformer) / 0.45 (First LGBM) / 0.15 (Second LGBM) ratio.\nAUC: 0.793 (slightly higher than the 1st Inference)\n\n# Training Environment\nOur biggest mistake was thinking that all the feature engineering, training, and inferencing process must be done in the Kaggle environment. We were only using Kaggle environment until 2 weeks before the competiton ended, and from then we started to use Google Colab with GPU and 25GB of RAM. Still, there were several times when Colab took GPU from us and didn't give it for several hours as we were constanly using their GPU.\n\nThanks to my teammates @shhrkre, @ysgong, @sanmaruum, @kuraji.",
      "votes": null
    },
    {
      "id": "1165470",
      "postDate": "01/23/2021 02:06:18",
      "content": "<p>good work. THX👍</p>",
      "rawMarkdown": "good work. THX👍",
      "votes": null
    },
    {
      "id": "1183116",
      "postDate": "02/02/2021 18:53:58",
      "content": "<p>nice sharing</p>",
      "rawMarkdown": "nice sharing",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1165470,
      "author_name": "shhrkre",
      "author_url": "",
      "post_date": "01/23/2021 02:06:18",
      "content": "<p>good work. THX👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1183116,
      "author_name": "henryhzy",
      "author_url": "",
      "post_date": "02/02/2021 18:53:58",
      "content": "<p>nice sharing</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1156923": "Also available at https://github.com/yoonseok312/riiid-answer-correctness-prediction with codes.\n\nThis is a solution write up for our model, which is an ensemble between a single Light Gradient Boosted Machine model, and a single Encoder-Decoder based Transformer model. It was our first time competing in a Kaggle competition and none of us had previous AI/ML/Stats experience so we learned a lot throughout this competition. Should you have any questions, feel free to contact me at yoonseok@berkeley.edu.\n\n# LightGBM\n## Features\nts_delta: the gap between timestamp of current content with previous content of the same user.\ntask_container_id\nprior_question_elapsed_time: in second, rounded\nprior_question_had_explanation\npart\nnum_tag: number of tags in that question\nu_chance: the average correctness of the user until current time\nu_attempts: number of content the user have done\nu_attempt_c: number of times the user interacted with the specific content in the past (only counting from >1 interactions due to memory)\nc_chance: the average correctness of the question until current time\nc_attempts: number of encounter of that question (all user)\nu_part_chance: the average correctness of the user doing the same part as the question\nu_part_attempts: number of question of the same part the user have done\nu_skill_chance: the average correctness of the user doing the same skill as the question (part < 5: listening, part >= 5: reading)\nu_skill_attempts: number of question of the same skill the user have done\nt_chance: the average correctness of the user of questions with specific tag until current time\nt_attempts: user's number of encounter of that tag\ntotal_explained: number of times explanation was provided to the user until current time (all contents)\n10_recent_correctness: user correctness of the most recent questions (up to 10)\n10_recent_mean_gap: mean ts_delta of the most recent questions (up to 10)\nbundle_elapsed: the mean elapsed time of the bundle, up until the abs time.\nmean_elapsed: the mean elapsed time of the user until now.\nprev_t1: tag of the last question\nprev_cor: correctness of the last question\ntrueskill_possibility: possibility of the user 'beating' the question (getting the question correct) based on trueskill\nmu: mu value (mean of trueskill ratings) of user\nsigma: sigma value (standard deviation of trueskill ratings) of user\nColumns with NaN value was filled with -1.\n\n## Cross Validation and train strategy\n1. Define an absolute time for the whole database (abs_time = user_id//50 + timestamp//1000)\n2. Sort by abs_time. All features mentioned above were engineered such that we will not take data from the future (higher abs_time) into account.\n3. Drop first 25% of the data. This data contains some noise, i.e. when nobody studied plenty of questions yet.\n4. Take last 25% of data as validation set.\n5. Train the model with remaining 50% of the data.\n\n## Single Model AUC\nWe used less than 30 features, but considering that most of the single LGBM models above 0.79 AUC used 40+ features, we did a decent work on focusing on imoportant features.\nLB score: AUC 0.789\nNumber of epochs: 6650 (around 15 hours of training in total)\n\n# Transformer\n## Encoder\nAdded below layers with positional encoding.\n\n1) Excercise Related\nmin_delta: minute difference from between this question and the previous. Cap at 1443 (1 day)\nday_delta: day difference from between this question and the previous. Cap at 30\nmonth_delta: month difference from between this question and the previous. Cap at 9\ntid: task container id\nis_with: if the question is presented with another question. Usually have the same task container\nc_part: part, one hot encode and denote skill (listening, reading, part1,2,...)\ntag1...6: tags of question (t1 to t6 are the tag of one question, t1 being the most important tag.) Above embeddings or Dense layer concatenated.\n\n2) Content id (cid)\nDense layer\n\n## Decoder\nAdded below layers with positional encoding.\n\n1) Response Related\nprev_answered_correct: correctness of previous answer.\nprior_elapsed: prior elapsed tiem\nprior_explained: prior has explanation Above embeddings or Dense layer concatenated.\n\n2) Answered Correctly\nConcatenate Lecture related Embeddings/Dense\nnum_lect: number of lecture the user have seen\nlec_type: 1 hot encode of most recent lecture type, (llecty1, 2...)\nlec_h_past: time since most recent lecture Above embeddings or Dense layer concatenated.\n\n## Parameters\nWINDOW_SIZE: 100\nEMBED_DIM: 256\nNUM_HEADS: 16\n\n## Cross Validation and train strategy\nUse first 80% of data as train set and last 20% as validation set.\n\n## Single Model AUC\nAUC 0.786\nSAINT model has plenty of room for improvement, but as 1 epoch took more than 10 hours to train we decided to focus on improving LGBM.\n\n# Inference\n## 1. Ensembling two models\nEnsembled a single LGBM model and a Transformer model in 0.55 (LGBM) / 0.45 ratio.\nAUC: 0.793\n\n## 2. Ensembling three models\nEnsembled two LGBM and a Transformer model. 2nd LGBM was same as the first LGBM but except features related to Trueskill. When 2 models out of three models predicted that the user is likely to answer correctly, we used the max value among the 3 predictions. When 2 models out of three models predicted that the user is likely to answer wronly, we used the min value among the 3 predictions. For remaining cases, we mixed three models in 0.4 (Transformer) / 0.45 (First LGBM) / 0.15 (Second LGBM) ratio.\nAUC: 0.793 (slightly higher than the 1st Inference)\n\n# Training Environment\nOur biggest mistake was thinking that all the feature engineering, training, and inferencing process must be done in the Kaggle environment. We were only using Kaggle environment until 2 weeks before the competiton ended, and from then we started to use Google Colab with GPU and 25GB of RAM. Still, there were several times when Colab took GPU from us and didn't give it for several hours as we were constanly using their GPU.\n\nThanks to my teammates @shhrkre, @ysgong, @sanmaruum, @kuraji.",
    "1165470": "good work. THX👍",
    "1183116": "nice sharing"
  },
  "source": "meta"
}