{
  "id": 209635,
  "title": "12th place Solution",
  "url": "/competitions/riiid-test-answer-prediction/writeups/pols-12th-place-solution",
  "author_name": "",
  "post_date": "2021-01-08T04:55:40.507274100Z",
  "votes": 74,
  "comment_count": 12,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F843848%2F94885ef1fcce200915129d79e234b4ea%2Friiid-solution.png?generation=1610081122634658&amp;alt=media\" alt=\"riiid-stack\"></p>\n<p>First of all, I want to thank both the hosts(kaggle and riiid) for hosting this competition.<br>\nIt was a really challenging competition in many ways, and really tested my limits. <br>\nThe competition was also very clean (albeit the privateLB bug unrelated to this competition). A very clean and logical train/test split/correlation is what makes this competition shine.</p>\n<hr>\n<h1>Overall solution:</h1>\n<p>For our single mode, NN scored around 0.806-810(privateLB), LGBM 0.804.<br>\nWe stacked these single models with NN and LGBM to achieve a score of 0.816.<br>\nI suspect our single models were mediocre in terms of score, but the stacking got us to gold.</p>\n<p>Regarding our 5place drop on the privateLB, we kind of expected it. We knew that one of our models(SAINT) had a bad private score because of the private score leak, but we couldn't quite fix it. Even after the competition has ended, we still have no clue as to why the model suffers in the privateLB.</p>\n<hr>\n<h1>Validation:</h1>\n<p>We used tito-split throughout most of the competition period.<br>\nWe mostly experienced good Valid-LB correlation, which I believe is the same for others. The single models used a 95M/5M split, and the stack used the 5M for training/validation.<br>\nWe also tried user-split as our final submission for diversity, but this had no impact on our final score.</p>\n<hr>\n<h1>LGBM details:</h1>\n<h4>Resource management</h4>\n<p>(When generating features for submission)<br>\nI used h5py for memory-hungry features like (user x content features, user x tag features).<br>\nThe h5py file uses user_id as the key, and read the whole user-feature when hitting new users while predicting.<br>\nThis creates a good balance between memory and runtime. Since there are not many unique users in the test-set (compared to train), the time to FileIO is not too much, and huge amounts of memory is saved.<br>\nAll the other features were pickled and loaded to memory in the beginning.</p>\n<h4>Features:</h4>\n<p>Question features:<br>\nThese were made by the full training-set(100M). <br>\nObviously this is leaky, but since every single question has &gt;1000 counts, the leak is tolerable.<br>\nExample features: mean question ac, mean user rating of who did not answer correctly.</p>\n<p>User features:<br>\nHow good the user is, especially related to parts, tags, contents.<br>\nExample features: user x contents mean question ac</p>\n<p>Timestamp features: <br>\nThis was kind of a surprise to me. Not only was the timestamp diff of t and t-1 a good features, diff of t-1 and t-2 up to t-9 and t-10 improved my model.<br>\nExample features: user user timestamp diff from last lecture</p>\n<p>Rating features:<br>\nElo features from this <a href=\"https://www.kaggle.com/stevemju/riiid-simple-elo-rating\" target=\"_blank\">notebook</a> (thank you very much). Trueskill did not improve my model.<br>\nWith only mean ac, the model cannot determine if the user is challenging hard questions or easy ones, so rating questions and users makes a lot of sense.</p>\n<p>SVD features:<br>\nLGBM is bad at expressing category columns (compared to NN). So I took the question embedding layer of NN, and used the 20dimension SVD as features.</p>\n<h4>Feature selection:</h4>\n<p>I used about 70-80 features for my model.<br>\nSince we were stacking a lot of models, I didn't want to use too much resource with my LGBM(runtime, memory), so I picked features which had a lot of impact, and made my model contribute to the stacking-model. There were a lot of features which didn't improve my model much, and all of them were thrown away. </p>\n<h4>Hyper-param:</h4>\n<p>I didn't change this too much, but increasing the num-leaf 127-&gt;1023 improved my score by 0.003, which was a surprise. This happened after I added lots of timestamp features, so there might be a very complicated interaction underlying in the timestamp.</p>\n<h4>Machine:</h4>\n<p>I used GCP, 64coreCPU 416GBmemory. Even with this monster machine, I ran out of memory a lot when generating the full features (which I avoided by processing in chunks).<br>\nI believe this instance cost me around $1000 over the competition (this is all payed by my company, and we are hiring btw). A lot of this cost happened because I was lazy (never used a preemptive instance), and I believe you could still be competitive with a $100 budget, even on this HUGE data, so don't be discouraged by the cost if you are a Kaggle beginner (I would recommend a competition with smaller data though :) )</p>\n<hr>\n<h1>Stacking details:</h1>\n<p>Stacking is always difficult. It often leads to overfitting.<br>\nIn this competition, there was a very good Valid-LB correlation, so I guessed(correctly) that stacking could work.<br>\nI wanted to make sure this layer works, so I tried to do everything conservative.</p>\n<h4>Validation:</h4>\n<p>3M/2M train/valid setup. For the final subs, I made multiple models with a time-series-split (negligible gain).</p>\n<h4>Input Models:</h4>\n<p>Multiple models from Sakami(SAINT based and AKT based), which had different window size. We couldn't fit in all the models, so we chose 3 as our final sub.<br>\nOwruby had another SAINT model which both added diversity and improved the stack a lot. Lyaka had a SSAKT model which also helped slightly.<br>\nI also added my LGBM model, which surprisingly improved the stack, even with the features added.</p>\n<h4>Input Features:</h4>\n<p>LGBM stack model: Hand selected 15 features from my single LGBM model. There was literally no gain from the other 65.<br>\nNN stack model: Selected features + the last layers from singleNNs. This was done by Lyaka.</p>\n<h4>Stacking models:</h4>\n<p>LGBM stack model: num_leaves was reduced 1023-&gt;127. <br>\nNN stack model: Lyaka did both MLP and a Transformer. Both had similar scores, and we used MLP as the final sub.</p>\n<h4>Final output:</h4>\n<p>Mean of LGBM and NN stack model</p>\n<hr>\n<h1>Final thoughts:</h1>\n<p>Although we couldn't quite achieve the goal we wanted (beating mamas), I think we tried our best and did our best.<br>\nEveryone contributed to the final stack, everyone did a ton of work, and we had great teamwork to make the stacking happen. I am really happy with the team we had, like I always had throughout my kaggle history. Thank you all.</p>",
  "messages": [
    {
      "id": "1143831",
      "postDate": "01/08/2021 04:55:40",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F843848%2F94885ef1fcce200915129d79e234b4ea%2Friiid-solution.png?generation=1610081122634658&amp;alt=media\" alt=\"riiid-stack\"></p>\n<p>First of all, I want to thank both the hosts(kaggle and riiid) for hosting this competition.<br>\nIt was a really challenging competition in many ways, and really tested my limits. <br>\nThe competition was also very clean (albeit the privateLB bug unrelated to this competition). A very clean and logical train/test split/correlation is what makes this competition shine.</p>\n<hr>\n<h1>Overall solution:</h1>\n<p>For our single mode, NN scored around 0.806-810(privateLB), LGBM 0.804.<br>\nWe stacked these single models with NN and LGBM to achieve a score of 0.816.<br>\nI suspect our single models were mediocre in terms of score, but the stacking got us to gold.</p>\n<p>Regarding our 5place drop on the privateLB, we kind of expected it. We knew that one of our models(SAINT) had a bad private score because of the private score leak, but we couldn't quite fix it. Even after the competition has ended, we still have no clue as to why the model suffers in the privateLB.</p>\n<hr>\n<h1>Validation:</h1>\n<p>We used tito-split throughout most of the competition period.<br>\nWe mostly experienced good Valid-LB correlation, which I believe is the same for others. The single models used a 95M/5M split, and the stack used the 5M for training/validation.<br>\nWe also tried user-split as our final submission for diversity, but this had no impact on our final score.</p>\n<hr>\n<h1>LGBM details:</h1>\n<h4>Resource management</h4>\n<p>(When generating features for submission)<br>\nI used h5py for memory-hungry features like (user x content features, user x tag features).<br>\nThe h5py file uses user_id as the key, and read the whole user-feature when hitting new users while predicting.<br>\nThis creates a good balance between memory and runtime. Since there are not many unique users in the test-set (compared to train), the time to FileIO is not too much, and huge amounts of memory is saved.<br>\nAll the other features were pickled and loaded to memory in the beginning.</p>\n<h4>Features:</h4>\n<p>Question features:<br>\nThese were made by the full training-set(100M). <br>\nObviously this is leaky, but since every single question has &gt;1000 counts, the leak is tolerable.<br>\nExample features: mean question ac, mean user rating of who did not answer correctly.</p>\n<p>User features:<br>\nHow good the user is, especially related to parts, tags, contents.<br>\nExample features: user x contents mean question ac</p>\n<p>Timestamp features: <br>\nThis was kind of a surprise to me. Not only was the timestamp diff of t and t-1 a good features, diff of t-1 and t-2 up to t-9 and t-10 improved my model.<br>\nExample features: user user timestamp diff from last lecture</p>\n<p>Rating features:<br>\nElo features from this <a href=\"https://www.kaggle.com/stevemju/riiid-simple-elo-rating\" target=\"_blank\">notebook</a> (thank you very much). Trueskill did not improve my model.<br>\nWith only mean ac, the model cannot determine if the user is challenging hard questions or easy ones, so rating questions and users makes a lot of sense.</p>\n<p>SVD features:<br>\nLGBM is bad at expressing category columns (compared to NN). So I took the question embedding layer of NN, and used the 20dimension SVD as features.</p>\n<h4>Feature selection:</h4>\n<p>I used about 70-80 features for my model.<br>\nSince we were stacking a lot of models, I didn't want to use too much resource with my LGBM(runtime, memory), so I picked features which had a lot of impact, and made my model contribute to the stacking-model. There were a lot of features which didn't improve my model much, and all of them were thrown away. </p>\n<h4>Hyper-param:</h4>\n<p>I didn't change this too much, but increasing the num-leaf 127-&gt;1023 improved my score by 0.003, which was a surprise. This happened after I added lots of timestamp features, so there might be a very complicated interaction underlying in the timestamp.</p>\n<h4>Machine:</h4>\n<p>I used GCP, 64coreCPU 416GBmemory. Even with this monster machine, I ran out of memory a lot when generating the full features (which I avoided by processing in chunks).<br>\nI believe this instance cost me around $1000 over the competition (this is all payed by my company, and we are hiring btw). A lot of this cost happened because I was lazy (never used a preemptive instance), and I believe you could still be competitive with a $100 budget, even on this HUGE data, so don't be discouraged by the cost if you are a Kaggle beginner (I would recommend a competition with smaller data though :) )</p>\n<hr>\n<h1>Stacking details:</h1>\n<p>Stacking is always difficult. It often leads to overfitting.<br>\nIn this competition, there was a very good Valid-LB correlation, so I guessed(correctly) that stacking could work.<br>\nI wanted to make sure this layer works, so I tried to do everything conservative.</p>\n<h4>Validation:</h4>\n<p>3M/2M train/valid setup. For the final subs, I made multiple models with a time-series-split (negligible gain).</p>\n<h4>Input Models:</h4>\n<p>Multiple models from Sakami(SAINT based and AKT based), which had different window size. We couldn't fit in all the models, so we chose 3 as our final sub.<br>\nOwruby had another SAINT model which both added diversity and improved the stack a lot. Lyaka had a SSAKT model which also helped slightly.<br>\nI also added my LGBM model, which surprisingly improved the stack, even with the features added.</p>\n<h4>Input Features:</h4>\n<p>LGBM stack model: Hand selected 15 features from my single LGBM model. There was literally no gain from the other 65.<br>\nNN stack model: Selected features + the last layers from singleNNs. This was done by Lyaka.</p>\n<h4>Stacking models:</h4>\n<p>LGBM stack model: num_leaves was reduced 1023-&gt;127. <br>\nNN stack model: Lyaka did both MLP and a Transformer. Both had similar scores, and we used MLP as the final sub.</p>\n<h4>Final output:</h4>\n<p>Mean of LGBM and NN stack model</p>\n<hr>\n<h1>Final thoughts:</h1>\n<p>Although we couldn't quite achieve the goal we wanted (beating mamas), I think we tried our best and did our best.<br>\nEveryone contributed to the final stack, everyone did a ton of work, and we had great teamwork to make the stacking happen. I am really happy with the team we had, like I always had throughout my kaggle history. Thank you all.</p>",
      "rawMarkdown": "![riiid-stack](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F843848%2F94885ef1fcce200915129d79e234b4ea%2Friiid-solution.png?generation=1610081122634658&alt=media)\n\nFirst of all, I want to thank both the hosts(kaggle and riiid) for hosting this competition.\nIt was a really challenging competition in many ways, and really tested my limits. \nThe competition was also very clean (albeit the privateLB bug unrelated to this competition). A very clean and logical train/test split/correlation is what makes this competition shine.\n\n---------------------------------------\n# Overall solution:\nFor our single mode, NN scored around 0.806-810(privateLB), LGBM 0.804.\nWe stacked these single models with NN and LGBM to achieve a score of 0.816.\nI suspect our single models were mediocre in terms of score, but the stacking got us to gold.\n\nRegarding our 5place drop on the privateLB, we kind of expected it. We knew that one of our models(SAINT) had a bad private score because of the private score leak, but we couldn't quite fix it. Even after the competition has ended, we still have no clue as to why the model suffers in the privateLB.\n\n---------------------------------------\n# Validation:\nWe used tito-split throughout most of the competition period.\nWe mostly experienced good Valid-LB correlation, which I believe is the same for others. The single models used a 95M/5M split, and the stack used the 5M for training/validation.\nWe also tried user-split as our final submission for diversity, but this had no impact on our final score.\n\n---------------------------------------\n# LGBM details:\n#### Resource management\n(When generating features for submission)\nI used h5py for memory-hungry features like (user x content features, user x tag features).\nThe h5py file uses user_id as the key, and read the whole user-feature when hitting new users while predicting.\nThis creates a good balance between memory and runtime. Since there are not many unique users in the test-set (compared to train), the time to FileIO is not too much, and huge amounts of memory is saved.\nAll the other features were pickled and loaded to memory in the beginning.\n\n#### Features: \nQuestion features:\nThese were made by the full training-set(100M). \nObviously this is leaky, but since every single question has >1000 counts, the leak is tolerable.\nExample features: mean question ac, mean user rating of who did not answer correctly.\n\nUser features:\nHow good the user is, especially related to parts, tags, contents.\nExample features: user x contents mean question ac\n\nTimestamp features: \nThis was kind of a surprise to me. Not only was the timestamp diff of t and t-1 a good features, diff of t-1 and t-2 up to t-9 and t-10 improved my model.\nExample features: user user timestamp diff from last lecture\n\nRating features:\nElo features from this [notebook](https://www.kaggle.com/stevemju/riiid-simple-elo-rating) (thank you very much). Trueskill did not improve my model.\nWith only mean ac, the model cannot determine if the user is challenging hard questions or easy ones, so rating questions and users makes a lot of sense.\n\nSVD features:\nLGBM is bad at expressing category columns (compared to NN). So I took the question embedding layer of NN, and used the 20dimension SVD as features.\n\n#### Feature selection:\nI used about 70-80 features for my model.\nSince we were stacking a lot of models, I didn't want to use too much resource with my LGBM(runtime, memory), so I picked features which had a lot of impact, and made my model contribute to the stacking-model. There were a lot of features which didn't improve my model much, and all of them were thrown away. \n\n#### Hyper-param:\nI didn't change this too much, but increasing the num-leaf 127->1023 improved my score by 0.003, which was a surprise. This happened after I added lots of timestamp features, so there might be a very complicated interaction underlying in the timestamp.\n\n#### Machine:\nI used GCP, 64coreCPU 416GBmemory. Even with this monster machine, I ran out of memory a lot when generating the full features (which I avoided by processing in chunks).\nI believe this instance cost me around $1000 over the competition (this is all payed by my company, and we are hiring btw). A lot of this cost happened because I was lazy (never used a preemptive instance), and I believe you could still be competitive with a $100 budget, even on this HUGE data, so don't be discouraged by the cost if you are a Kaggle beginner (I would recommend a competition with smaller data though :) )\n\n---------------------------------------\n# Stacking details:\nStacking is always difficult. It often leads to overfitting.\nIn this competition, there was a very good Valid-LB correlation, so I guessed(correctly) that stacking could work.\nI wanted to make sure this layer works, so I tried to do everything conservative.\n\n#### Validation:\n3M/2M train/valid setup. For the final subs, I made multiple models with a time-series-split (negligible gain).\n\n#### Input Models:\nMultiple models from Sakami(SAINT based and AKT based), which had different window size. We couldn't fit in all the models, so we chose 3 as our final sub.\nOwruby had another SAINT model which both added diversity and improved the stack a lot. Lyaka had a SSAKT model which also helped slightly.\nI also added my LGBM model, which surprisingly improved the stack, even with the features added.\n\n#### Input Features: \nLGBM stack model: Hand selected 15 features from my single LGBM model. There was literally no gain from the other 65.\nNN stack model: Selected features + the last layers from singleNNs. This was done by Lyaka.\n\n#### Stacking models:\nLGBM stack model: num_leaves was reduced 1023->127. \nNN stack model: Lyaka did both MLP and a Transformer. Both had similar scores, and we used MLP as the final sub.\n\n#### Final output:\nMean of LGBM and NN stack model\n\n---------------------------------------\n# Final thoughts:\nAlthough we couldn't quite achieve the goal we wanted (beating mamas), I think we tried our best and did our best.\nEveryone contributed to the final stack, everyone did a ton of work, and we had great teamwork to make the stacking happen. I am really happy with the team we had, like I always had throughout my kaggle history. Thank you all.",
      "votes": null
    },
    {
      "id": "1143837",
      "postDate": "01/08/2021 05:02:32",
      "content": "<p>NN details will be posted later by <a href=\"https://www.kaggle.com/sakami\" target=\"_blank\">@sakami</a> </p>",
      "rawMarkdown": "NN details will be posted later by @sakami",
      "votes": null
    },
    {
      "id": "1143859",
      "postDate": "01/08/2021 05:21:46",
      "content": "<p><a href=\"https://www.kaggle.com/pocket\" target=\"_blank\">@pocket</a> i would say its a great work . We did try AKT but it dknt gave us score ,we replaced mha with akt mha ,as part of change  not sure if this was right implementation </p>",
      "rawMarkdown": "pocket i would say its a great work . We did try AKT but it dknt gave us score ,we replaced mha with akt mha ,as part of change  not sure if this was right implementation",
      "votes": null
    },
    {
      "id": "1143998",
      "postDate": "01/08/2021 07:09:56",
      "content": "<p>Good job! I am curious, for people that stacked multiple transformers, how come your solution didn't time out. Mine did for sequences longer than 100, and it is a single 128 d_model. This discouraged me from scaling the model.</p>",
      "rawMarkdown": "Good job! I am curious, for people that stacked multiple transformers, how come your solution didn't time out. Mine did for sequences longer than 100, and it is a single 128 d_model. This discouraged me from scaling the model.",
      "votes": null
    },
    {
      "id": "1144010",
      "postDate": "01/08/2021 07:17:30",
      "content": "<p><a href=\"https://www.kaggle.com/pocketsuteado\" target=\"_blank\">@pocketsuteado</a> Can you share yours AKT net code plz?  we want to learn yours implement, thank you</p>",
      "rawMarkdown": "pocketsuteado Can you share yours AKT net code plz?  we want to learn yours implement, thank you",
      "votes": null
    },
    {
      "id": "1144064",
      "postDate": "01/08/2021 08:07:43",
      "content": "<p>Thank you for the detailed explanation of your solution. I believe that a great result are derived from combined efforts of all the team members. Congrats for gold medal, and becoming competition grandmaster <a href=\"https://www.kaggle.com/pocketsuteado\" target=\"_blank\">@pocketsuteado</a> and <a href=\"https://www.kaggle.com/owruby\" target=\"_blank\">@owruby</a> !</p>",
      "rawMarkdown": "Thank you for the detailed explanation of your solution. I believe that a great result are derived from combined efforts of all the team members. Congrats for gold medal, and becoming competition grandmaster @pocketsuteado and @owruby !",
      "votes": null
    },
    {
      "id": "1144097",
      "postDate": "01/08/2021 08:27:33",
      "content": "<p>You can check our SAINT/AKT code <a href=\"https://github.com/sakami0000/kaggle_riiid\" target=\"_blank\">here</a>.</p>",
      "rawMarkdown": "You can check our SAINT/AKT code [here](https://github.com/sakami0000/kaggle_riiid).",
      "votes": null
    },
    {
      "id": "1144099",
      "postDate": "01/08/2021 08:28:13",
      "content": "<p>Additional resources:<br>\n<a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209683\" target=\"_blank\">NN solution</a><br>\n<a href=\"https://github.com/backonhighway/gm_riiid\" target=\"_blank\">My code</a> (As is, and unorganized)<br>\n<a href=\"https://github.com/sakami0000/kaggle_riiid\" target=\"_blank\">NN code</a></p>",
      "rawMarkdown": "Additional resources:\n[NN solution](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209683)\n[My code](https://github.com/backonhighway/gm_riiid) (As is, and unorganized)\n[NN code](https://github.com/sakami0000/kaggle_riiid)",
      "votes": null
    },
    {
      "id": "1144101",
      "postDate": "01/08/2021 08:29:14",
      "content": "<p>You can check <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209683\" target=\"_blank\">our NN solution</a>, or ask <a href=\"https://www.kaggle.com/sakami\" target=\"_blank\">@sakami</a> </p>",
      "rawMarkdown": "You can check [our NN solution](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209683), or ask @sakami",
      "votes": null
    },
    {
      "id": "1144107",
      "postDate": "01/08/2021 08:32:25",
      "content": "<p>Nice, thank you</p>",
      "rawMarkdown": "Nice, thank you",
      "votes": null
    },
    {
      "id": "1144372",
      "postDate": "01/08/2021 12:28:09",
      "content": "<p>Thanks for the nice write-up (I always enjoy pretty stacking schemes) and congratz on the nice finish ! </p>",
      "rawMarkdown": "Thanks for the nice write-up (I always enjoy pretty stacking schemes) and congratz on the nice finish !",
      "votes": null
    },
    {
      "id": "1145139",
      "postDate": "01/08/2021 23:36:23",
      "content": "<p>Thanks for sharing!  And huge congrats on becoming GM! </p>",
      "rawMarkdown": "Thanks for sharing!  And huge congrats on becoming GM!",
      "votes": null
    },
    {
      "id": "1153926",
      "postDate": "01/15/2021 08:53:11",
      "content": "<p>Thanks for sharing! Very inspirational</p>",
      "rawMarkdown": "Thanks for sharing! Very inspirational",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1143837,
      "author_name": "pocketsuteado",
      "author_url": "",
      "post_date": "01/08/2021 05:02:32",
      "content": "<p>NN details will be posted later by <a href=\"https://www.kaggle.com/sakami\" target=\"_blank\">@sakami</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1143859,
      "author_name": "jaideepvalani",
      "author_url": "",
      "post_date": "01/08/2021 05:21:46",
      "content": "<p><a href=\"https://www.kaggle.com/pocket\" target=\"_blank\">@pocket</a> i would say its a great work . We did try AKT but it dknt gave us score ,we replaced mha with akt mha ,as part of change  not sure if this was right implementation </p>",
      "votes": null,
      "replies": [
        {
          "id": 1144010,
          "author_name": "cswwp347724",
          "author_url": "",
          "post_date": "01/08/2021 07:17:30",
          "content": "<p><a href=\"https://www.kaggle.com/pocketsuteado\" target=\"_blank\">@pocketsuteado</a> Can you share yours AKT net code plz?  we want to learn yours implement, thank you</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1144097,
          "author_name": "sakami",
          "author_url": "",
          "post_date": "01/08/2021 08:27:33",
          "content": "<p>You can check our SAINT/AKT code <a href=\"https://github.com/sakami0000/kaggle_riiid\" target=\"_blank\">here</a>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1144107,
          "author_name": "cswwp347724",
          "author_url": "",
          "post_date": "01/08/2021 08:32:25",
          "content": "<p>Nice, thank you</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1143998,
      "author_name": "abdessalemboukil",
      "author_url": "",
      "post_date": "01/08/2021 07:09:56",
      "content": "<p>Good job! I am curious, for people that stacked multiple transformers, how come your solution didn't time out. Mine did for sequences longer than 100, and it is a single 128 d_model. This discouraged me from scaling the model.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1144101,
          "author_name": "pocketsuteado",
          "author_url": "",
          "post_date": "01/08/2021 08:29:14",
          "content": "<p>You can check <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209683\" target=\"_blank\">our NN solution</a>, or ask <a href=\"https://www.kaggle.com/sakami\" target=\"_blank\">@sakami</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1144064,
      "author_name": "sishihara",
      "author_url": "",
      "post_date": "01/08/2021 08:07:43",
      "content": "<p>Thank you for the detailed explanation of your solution. I believe that a great result are derived from combined efforts of all the team members. Congrats for gold medal, and becoming competition grandmaster <a href=\"https://www.kaggle.com/pocketsuteado\" target=\"_blank\">@pocketsuteado</a> and <a href=\"https://www.kaggle.com/owruby\" target=\"_blank\">@owruby</a> !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1144099,
      "author_name": "pocketsuteado",
      "author_url": "",
      "post_date": "01/08/2021 08:28:13",
      "content": "<p>Additional resources:<br>\n<a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209683\" target=\"_blank\">NN solution</a><br>\n<a href=\"https://github.com/backonhighway/gm_riiid\" target=\"_blank\">My code</a> (As is, and unorganized)<br>\n<a href=\"https://github.com/sakami0000/kaggle_riiid\" target=\"_blank\">NN code</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1144372,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "01/08/2021 12:28:09",
      "content": "<p>Thanks for the nice write-up (I always enjoy pretty stacking schemes) and congratz on the nice finish ! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1145139,
      "author_name": "raphael1123",
      "author_url": "",
      "post_date": "01/08/2021 23:36:23",
      "content": "<p>Thanks for sharing!  And huge congrats on becoming GM! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1153926,
      "author_name": "neilgibbons",
      "author_url": "",
      "post_date": "01/15/2021 08:53:11",
      "content": "<p>Thanks for sharing! Very inspirational</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1143831": "![riiid-stack](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F843848%2F94885ef1fcce200915129d79e234b4ea%2Friiid-solution.png?generation=1610081122634658&alt=media)\n\nFirst of all, I want to thank both the hosts(kaggle and riiid) for hosting this competition.\nIt was a really challenging competition in many ways, and really tested my limits. \nThe competition was also very clean (albeit the privateLB bug unrelated to this competition). A very clean and logical train/test split/correlation is what makes this competition shine.\n\n---------------------------------------\n# Overall solution:\nFor our single mode, NN scored around 0.806-810(privateLB), LGBM 0.804.\nWe stacked these single models with NN and LGBM to achieve a score of 0.816.\nI suspect our single models were mediocre in terms of score, but the stacking got us to gold.\n\nRegarding our 5place drop on the privateLB, we kind of expected it. We knew that one of our models(SAINT) had a bad private score because of the private score leak, but we couldn't quite fix it. Even after the competition has ended, we still have no clue as to why the model suffers in the privateLB.\n\n---------------------------------------\n# Validation:\nWe used tito-split throughout most of the competition period.\nWe mostly experienced good Valid-LB correlation, which I believe is the same for others. The single models used a 95M/5M split, and the stack used the 5M for training/validation.\nWe also tried user-split as our final submission for diversity, but this had no impact on our final score.\n\n---------------------------------------\n# LGBM details:\n#### Resource management\n(When generating features for submission)\nI used h5py for memory-hungry features like (user x content features, user x tag features).\nThe h5py file uses user_id as the key, and read the whole user-feature when hitting new users while predicting.\nThis creates a good balance between memory and runtime. Since there are not many unique users in the test-set (compared to train), the time to FileIO is not too much, and huge amounts of memory is saved.\nAll the other features were pickled and loaded to memory in the beginning.\n\n#### Features: \nQuestion features:\nThese were made by the full training-set(100M). \nObviously this is leaky, but since every single question has >1000 counts, the leak is tolerable.\nExample features: mean question ac, mean user rating of who did not answer correctly.\n\nUser features:\nHow good the user is, especially related to parts, tags, contents.\nExample features: user x contents mean question ac\n\nTimestamp features: \nThis was kind of a surprise to me. Not only was the timestamp diff of t and t-1 a good features, diff of t-1 and t-2 up to t-9 and t-10 improved my model.\nExample features: user user timestamp diff from last lecture\n\nRating features:\nElo features from this [notebook](https://www.kaggle.com/stevemju/riiid-simple-elo-rating) (thank you very much). Trueskill did not improve my model.\nWith only mean ac, the model cannot determine if the user is challenging hard questions or easy ones, so rating questions and users makes a lot of sense.\n\nSVD features:\nLGBM is bad at expressing category columns (compared to NN). So I took the question embedding layer of NN, and used the 20dimension SVD as features.\n\n#### Feature selection:\nI used about 70-80 features for my model.\nSince we were stacking a lot of models, I didn't want to use too much resource with my LGBM(runtime, memory), so I picked features which had a lot of impact, and made my model contribute to the stacking-model. There were a lot of features which didn't improve my model much, and all of them were thrown away. \n\n#### Hyper-param:\nI didn't change this too much, but increasing the num-leaf 127->1023 improved my score by 0.003, which was a surprise. This happened after I added lots of timestamp features, so there might be a very complicated interaction underlying in the timestamp.\n\n#### Machine:\nI used GCP, 64coreCPU 416GBmemory. Even with this monster machine, I ran out of memory a lot when generating the full features (which I avoided by processing in chunks).\nI believe this instance cost me around $1000 over the competition (this is all payed by my company, and we are hiring btw). A lot of this cost happened because I was lazy (never used a preemptive instance), and I believe you could still be competitive with a $100 budget, even on this HUGE data, so don't be discouraged by the cost if you are a Kaggle beginner (I would recommend a competition with smaller data though :) )\n\n---------------------------------------\n# Stacking details:\nStacking is always difficult. It often leads to overfitting.\nIn this competition, there was a very good Valid-LB correlation, so I guessed(correctly) that stacking could work.\nI wanted to make sure this layer works, so I tried to do everything conservative.\n\n#### Validation:\n3M/2M train/valid setup. For the final subs, I made multiple models with a time-series-split (negligible gain).\n\n#### Input Models:\nMultiple models from Sakami(SAINT based and AKT based), which had different window size. We couldn't fit in all the models, so we chose 3 as our final sub.\nOwruby had another SAINT model which both added diversity and improved the stack a lot. Lyaka had a SSAKT model which also helped slightly.\nI also added my LGBM model, which surprisingly improved the stack, even with the features added.\n\n#### Input Features: \nLGBM stack model: Hand selected 15 features from my single LGBM model. There was literally no gain from the other 65.\nNN stack model: Selected features + the last layers from singleNNs. This was done by Lyaka.\n\n#### Stacking models:\nLGBM stack model: num_leaves was reduced 1023->127. \nNN stack model: Lyaka did both MLP and a Transformer. Both had similar scores, and we used MLP as the final sub.\n\n#### Final output:\nMean of LGBM and NN stack model\n\n---------------------------------------\n# Final thoughts:\nAlthough we couldn't quite achieve the goal we wanted (beating mamas), I think we tried our best and did our best.\nEveryone contributed to the final stack, everyone did a ton of work, and we had great teamwork to make the stacking happen. I am really happy with the team we had, like I always had throughout my kaggle history. Thank you all.",
    "1143837": "NN details will be posted later by @sakami",
    "1143859": "pocket i would say its a great work . We did try AKT but it dknt gave us score ,we replaced mha with akt mha ,as part of change  not sure if this was right implementation",
    "1143998": "Good job! I am curious, for people that stacked multiple transformers, how come your solution didn't time out. Mine did for sequences longer than 100, and it is a single 128 d_model. This discouraged me from scaling the model.",
    "1144010": "pocketsuteado Can you share yours AKT net code plz?  we want to learn yours implement, thank you",
    "1144064": "Thank you for the detailed explanation of your solution. I believe that a great result are derived from combined efforts of all the team members. Congrats for gold medal, and becoming competition grandmaster @pocketsuteado and @owruby !",
    "1144097": "You can check our SAINT/AKT code [here](https://github.com/sakami0000/kaggle_riiid).",
    "1144099": "Additional resources:\n[NN solution](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209683)\n[My code](https://github.com/backonhighway/gm_riiid) (As is, and unorganized)\n[NN code](https://github.com/sakami0000/kaggle_riiid)",
    "1144101": "You can check [our NN solution](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209683), or ask @sakami",
    "1144107": "Nice, thank you",
    "1144372": "Thanks for the nice write-up (I always enjoy pretty stacking schemes) and congratz on the nice finish !",
    "1145139": "Thanks for sharing!  And huge congrats on becoming GM!",
    "1153926": "Thanks for sharing! Very inspirational"
  },
  "source": "meta"
}