{
  "id": 209619,
  "title": "158th Solution LGBM + FTRL, Private LB: 0793",
  "url": "/competitions/riiid-test-answer-prediction/writeups/158th-solution-lgbm-ftrl-private-lb-0793",
  "author_name": "",
  "post_date": "2021-01-08T03:20:09.043204800Z",
  "votes": 18,
  "comment_count": 2,
  "views": 0,
  "content": "<h2>Huge thanks to:</h2>\n<p>First and foremost I would like to thank the coordinators of RIIID for conducting this awesome competition and all the Kaggle Staffs who kept their constant vigilance for bugs and issues. Despite having data leaks and many such issues, the kaggle team handled the issue promptly and efficiently.  The competition API was one of the most frustrating and yet necessary aspect of a competition such as this one. </p>\n<p>The list of people I would like to thank is numerous and large, a few notable mentions but not limited to are:</p>\n<ol>\n<li><a href=\"https://www.kaggle.com/shoheiazuma/riiid-lgbm-starter:\" target=\"_blank\">https://www.kaggle.com/shoheiazuma/riiid-lgbm-starter:</a> The dataset was huge and intimidating, <a href=\"https://www.kaggle.com/shoheiazuma\" target=\"_blank\">@shoheiazuma</a> notebook gave the idea to train on just the last 24 records. I improvised on this idea a lot. Using final 30M records for training felt like I was missing out some users. Training on final 250 records of each user gave me comparable results to training on the entire dataset. </li>\n<li><a href=\"https://www.kaggle.com/vopani\" target=\"_blank\">@vopani</a>'s classic notebooks: <a href=\"https://www.kaggle.com/rohanrao/tutorial-on-reading-large-datasets\" target=\"_blank\">this</a> and <a href=\"https://www.kaggle.com/rohanrao/riiid-ftrl-ftw\" target=\"_blank\">this</a>: These tutorials Introduced me to my now favorite feather format. It also introduced me to DataTable FTRL and Dask. </li>\n<li><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> discussions on repeated user questions and his ideas to use Bit Arrays. He contributed a ton of ideas and intrigued me to come with cleverer features. </li>\n<li><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/196009:\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/196009:</a> Ditch pandas! All my user state updations were done in a loop. Easier on the pipeline and easier for my brain too :)</li>\n<li><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/196942:\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/196942:</a> A kaggle GM <a href=\"https://www.kaggle.com/silogram\" target=\"_blank\">@silogram</a> using LGBM to hit high ranks inspired me not to give up so soon when my score stuck at local optimas.</li>\n<li><a href=\"https://www.kaggle.com/anuragtr\" target=\"_blank\">@anuragtr</a> discussion <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/198245\" target=\"_blank\">here</a> on how to prevent ram spikes when starting LGB training.</li>\n</ol>\n<h2>The Model Idea and difficulties faced:</h2>\n<p>So what did <em>I</em> do? My model idea is quite simple: <strong>8 bagged LGBM models fed into an FTRL model</strong>. Kaggle CPU had memory restrictions of 16GB. Even concatenating all the features together threw that horrible OOM error. I created features piece by piece and concatenated them into three different chunks to create the training data. </p>\n<p>Now we have 100M data containing 38 Features. Training an LGB on this large data would still throw an OOM. So I decided to do bagging (which would still throw an OOM error if you didn't slowly convert the data into float numpy arrays yourself using loops). Now Bagging is not an option you'd go for if you wish to make more accurate predictions, but a discussion mentioned that training on larger dataset gave a score boost of ~0.01 and so I did. </p>\n<p>Training on the first chunk (25M) gave a CV of 0.791. After 7 such chunks, I got a CV of 0.792. Frustrated by seeing my LB ranks drop and to improve the predictive power a bit, I decided to add another model to the mix, a FTRL. </p>\n<p>I was also motivated to use FTRL for two reasons: </p>\n<ol>\n<li>I only had two days left before competition ends &amp; FTRL was super fast!</li>\n<li>FTRL model is best suited for applications were the predictions need to dynamic and evolving. </li>\n</ol>\n<p>FTRL by itself didn't perform very well on my CVs. So I decided to feed my LGB model predictions to it to help it make its predictions. If the FTRL sees the LGB making false predictions during inference, it can improve or alter its predictions accordingly. Doing this improved my LB from 0.792 to 0.793. I have no prior knowledge with FTRL models. Perhaps with better features FTRL might have worked even better.</p>\n<h2>Some of my personal takeaways:</h2>\n<ul>\n<li><em>Don't shy away from newer unfamiliar concepts!</em> With such a huge time frame given, I could have learned and implemented a transformer model. Instead I chose to stick with what I knew best: Brute forcing my way with hand coded features. Although this approach taught me how to make the most out of the limited given resources.</li>\n<li><em>Numba is awesome!</em> Ton of code speed up, all with a simple decorator! Numba although fast, may throw OOM if blindly used. My notebook (soon to be published) demonstrates how I leveraged this awesome library.</li>\n<li><em>LB ranks doesn't matter!</em> To be honest, It was painful to see my LB ranks drop from 60 all the way down to 158. But in the end, I managed to accept it and even laugh about it! See <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/208356\" target=\"_blank\">this</a> hilarious meme thread.</li>\n</ul>\n<hr>\n<p>I enjoyed participating in this competition. It gave me so much more than I imagined it would. To all the winners and participants of this competition: <em>My heartiest congratulations and a Happy new year!</em></p>\n<hr>",
  "messages": [
    {
      "id": "1143738",
      "postDate": "01/08/2021 03:20:09",
      "content": "<h2>Huge thanks to:</h2>\n<p>First and foremost I would like to thank the coordinators of RIIID for conducting this awesome competition and all the Kaggle Staffs who kept their constant vigilance for bugs and issues. Despite having data leaks and many such issues, the kaggle team handled the issue promptly and efficiently.  The competition API was one of the most frustrating and yet necessary aspect of a competition such as this one. </p>\n<p>The list of people I would like to thank is numerous and large, a few notable mentions but not limited to are:</p>\n<ol>\n<li><a href=\"https://www.kaggle.com/shoheiazuma/riiid-lgbm-starter:\" target=\"_blank\">https://www.kaggle.com/shoheiazuma/riiid-lgbm-starter:</a> The dataset was huge and intimidating, <a href=\"https://www.kaggle.com/shoheiazuma\" target=\"_blank\">@shoheiazuma</a> notebook gave the idea to train on just the last 24 records. I improvised on this idea a lot. Using final 30M records for training felt like I was missing out some users. Training on final 250 records of each user gave me comparable results to training on the entire dataset. </li>\n<li><a href=\"https://www.kaggle.com/vopani\" target=\"_blank\">@vopani</a>'s classic notebooks: <a href=\"https://www.kaggle.com/rohanrao/tutorial-on-reading-large-datasets\" target=\"_blank\">this</a> and <a href=\"https://www.kaggle.com/rohanrao/riiid-ftrl-ftw\" target=\"_blank\">this</a>: These tutorials Introduced me to my now favorite feather format. It also introduced me to DataTable FTRL and Dask. </li>\n<li><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> discussions on repeated user questions and his ideas to use Bit Arrays. He contributed a ton of ideas and intrigued me to come with cleverer features. </li>\n<li><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/196009:\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/196009:</a> Ditch pandas! All my user state updations were done in a loop. Easier on the pipeline and easier for my brain too :)</li>\n<li><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/196942:\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/196942:</a> A kaggle GM <a href=\"https://www.kaggle.com/silogram\" target=\"_blank\">@silogram</a> using LGBM to hit high ranks inspired me not to give up so soon when my score stuck at local optimas.</li>\n<li><a href=\"https://www.kaggle.com/anuragtr\" target=\"_blank\">@anuragtr</a> discussion <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/198245\" target=\"_blank\">here</a> on how to prevent ram spikes when starting LGB training.</li>\n</ol>\n<h2>The Model Idea and difficulties faced:</h2>\n<p>So what did <em>I</em> do? My model idea is quite simple: <strong>8 bagged LGBM models fed into an FTRL model</strong>. Kaggle CPU had memory restrictions of 16GB. Even concatenating all the features together threw that horrible OOM error. I created features piece by piece and concatenated them into three different chunks to create the training data. </p>\n<p>Now we have 100M data containing 38 Features. Training an LGB on this large data would still throw an OOM. So I decided to do bagging (which would still throw an OOM error if you didn't slowly convert the data into float numpy arrays yourself using loops). Now Bagging is not an option you'd go for if you wish to make more accurate predictions, but a discussion mentioned that training on larger dataset gave a score boost of ~0.01 and so I did. </p>\n<p>Training on the first chunk (25M) gave a CV of 0.791. After 7 such chunks, I got a CV of 0.792. Frustrated by seeing my LB ranks drop and to improve the predictive power a bit, I decided to add another model to the mix, a FTRL. </p>\n<p>I was also motivated to use FTRL for two reasons: </p>\n<ol>\n<li>I only had two days left before competition ends &amp; FTRL was super fast!</li>\n<li>FTRL model is best suited for applications were the predictions need to dynamic and evolving. </li>\n</ol>\n<p>FTRL by itself didn't perform very well on my CVs. So I decided to feed my LGB model predictions to it to help it make its predictions. If the FTRL sees the LGB making false predictions during inference, it can improve or alter its predictions accordingly. Doing this improved my LB from 0.792 to 0.793. I have no prior knowledge with FTRL models. Perhaps with better features FTRL might have worked even better.</p>\n<h2>Some of my personal takeaways:</h2>\n<ul>\n<li><em>Don't shy away from newer unfamiliar concepts!</em> With such a huge time frame given, I could have learned and implemented a transformer model. Instead I chose to stick with what I knew best: Brute forcing my way with hand coded features. Although this approach taught me how to make the most out of the limited given resources.</li>\n<li><em>Numba is awesome!</em> Ton of code speed up, all with a simple decorator! Numba although fast, may throw OOM if blindly used. My notebook (soon to be published) demonstrates how I leveraged this awesome library.</li>\n<li><em>LB ranks doesn't matter!</em> To be honest, It was painful to see my LB ranks drop from 60 all the way down to 158. But in the end, I managed to accept it and even laugh about it! See <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/208356\" target=\"_blank\">this</a> hilarious meme thread.</li>\n</ul>\n<hr>\n<p>I enjoyed participating in this competition. It gave me so much more than I imagined it would. To all the winners and participants of this competition: <em>My heartiest congratulations and a Happy new year!</em></p>\n<hr>",
      "rawMarkdown": "## Huge thanks to:\nFirst and foremost I would like to thank the coordinators of RIIID for conducting this awesome competition and all the Kaggle Staffs who kept their constant vigilance for bugs and issues. Despite having data leaks and many such issues, the kaggle team handled the issue promptly and efficiently.  The competition API was one of the most frustrating and yet necessary aspect of a competition such as this one. \n\nThe list of people I would like to thank is numerous and large, a few notable mentions but not limited to are:\n1. https://www.kaggle.com/shoheiazuma/riiid-lgbm-starter: The dataset was huge and intimidating, @shoheiazuma notebook gave the idea to train on just the last 24 records. I improvised on this idea a lot. Using final 30M records for training felt like I was missing out some users. Training on final 250 records of each user gave me comparable results to training on the entire dataset. \n2. @vopani's classic notebooks: [this](https://www.kaggle.com/rohanrao/tutorial-on-reading-large-datasets) and [this](https://www.kaggle.com/rohanrao/riiid-ftrl-ftw): These tutorials Introduced me to my now favorite feather format. It also introduced me to DataTable FTRL and Dask. \n3. @adityaecdrid discussions on repeated user questions and his ideas to use Bit Arrays. He contributed a ton of ideas and intrigued me to come with cleverer features. \n4. https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/196009: Ditch pandas! All my user state updations were done in a loop. Easier on the pipeline and easier for my brain too :)\n5. https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/196942: A kaggle GM @silogram using LGBM to hit high ranks inspired me not to give up so soon when my score stuck at local optimas.\n6. @anuragtr discussion [here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/198245) on how to prevent ram spikes when starting LGB training.\n\n## The Model Idea and difficulties faced:\nSo what did *I* do? My model idea is quite simple: **8 bagged LGBM models fed into an FTRL model**. Kaggle CPU had memory restrictions of 16GB. Even concatenating all the features together threw that horrible OOM error. I created features piece by piece and concatenated them into three different chunks to create the training data. \n\nNow we have 100M data containing 38 Features. Training an LGB on this large data would still throw an OOM. So I decided to do bagging (which would still throw an OOM error if you didn't slowly convert the data into float numpy arrays yourself using loops). Now Bagging is not an option you'd go for if you wish to make more accurate predictions, but a discussion mentioned that training on larger dataset gave a score boost of ~0.01 and so I did. \n\nTraining on the first chunk (25M) gave a CV of 0.791. After 7 such chunks, I got a CV of 0.792. Frustrated by seeing my LB ranks drop and to improve the predictive power a bit, I decided to add another model to the mix, a FTRL. \n\nI was also motivated to use FTRL for two reasons: \n\n1. I only had two days left before competition ends & FTRL was super fast!\n2. FTRL model is best suited for applications were the predictions need to dynamic and evolving. \n\nFTRL by itself didn't perform very well on my CVs. So I decided to feed my LGB model predictions to it to help it make its predictions. If the FTRL sees the LGB making false predictions during inference, it can improve or alter its predictions accordingly. Doing this improved my LB from 0.792 to 0.793. I have no prior knowledge with FTRL models. Perhaps with better features FTRL might have worked even better.\n\n## Some of my personal takeaways:\n\n- *Don't shy away from newer unfamiliar concepts!* With such a huge time frame given, I could have learned and implemented a transformer model. Instead I chose to stick with what I knew best: Brute forcing my way with hand coded features. Although this approach taught me how to make the most out of the limited given resources.\n- *Numba is awesome!* Ton of code speed up, all with a simple decorator! Numba although fast, may throw OOM if blindly used. My notebook (soon to be published) demonstrates how I leveraged this awesome library.\n- *LB ranks doesn't matter!* To be honest, It was painful to see my LB ranks drop from 60 all the way down to 158. But in the end, I managed to accept it and even laugh about it! See [this](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/208356) hilarious meme thread.\n\n***\nI enjoyed participating in this competition. It gave me so much more than I imagined it would. To all the winners and participants of this competition: *My heartiest congratulations and a Happy new year!*\n***",
      "votes": null
    },
    {
      "id": "1143863",
      "postDate": "01/08/2021 05:31:51",
      "content": "<p>Nice Idea to use FTRLs!</p>",
      "rawMarkdown": "Nice Idea to use FTRLs!",
      "votes": null
    },
    {
      "id": "1144381",
      "postDate": "01/08/2021 12:36:46",
      "content": "<p>My notebook is now public. You can check them out <a href=\"https://www.kaggle.com/doctorkael/riiid-lgb-ftrl-data-gen-and-training-logic\" target=\"_blank\">here</a>!</p>",
      "rawMarkdown": "My notebook is now public. You can check them out [here](https://www.kaggle.com/doctorkael/riiid-lgb-ftrl-data-gen-and-training-logic)!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1143863,
      "author_name": "adityaecdrid",
      "author_url": "",
      "post_date": "01/08/2021 05:31:51",
      "content": "<p>Nice Idea to use FTRLs!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1144381,
      "author_name": "doctorkael",
      "author_url": "",
      "post_date": "01/08/2021 12:36:46",
      "content": "<p>My notebook is now public. You can check them out <a href=\"https://www.kaggle.com/doctorkael/riiid-lgb-ftrl-data-gen-and-training-logic\" target=\"_blank\">here</a>!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1143738": "## Huge thanks to:\nFirst and foremost I would like to thank the coordinators of RIIID for conducting this awesome competition and all the Kaggle Staffs who kept their constant vigilance for bugs and issues. Despite having data leaks and many such issues, the kaggle team handled the issue promptly and efficiently.  The competition API was one of the most frustrating and yet necessary aspect of a competition such as this one. \n\nThe list of people I would like to thank is numerous and large, a few notable mentions but not limited to are:\n1. https://www.kaggle.com/shoheiazuma/riiid-lgbm-starter: The dataset was huge and intimidating, @shoheiazuma notebook gave the idea to train on just the last 24 records. I improvised on this idea a lot. Using final 30M records for training felt like I was missing out some users. Training on final 250 records of each user gave me comparable results to training on the entire dataset. \n2. @vopani's classic notebooks: [this](https://www.kaggle.com/rohanrao/tutorial-on-reading-large-datasets) and [this](https://www.kaggle.com/rohanrao/riiid-ftrl-ftw): These tutorials Introduced me to my now favorite feather format. It also introduced me to DataTable FTRL and Dask. \n3. @adityaecdrid discussions on repeated user questions and his ideas to use Bit Arrays. He contributed a ton of ideas and intrigued me to come with cleverer features. \n4. https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/196009: Ditch pandas! All my user state updations were done in a loop. Easier on the pipeline and easier for my brain too :)\n5. https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/196942: A kaggle GM @silogram using LGBM to hit high ranks inspired me not to give up so soon when my score stuck at local optimas.\n6. @anuragtr discussion [here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/198245) on how to prevent ram spikes when starting LGB training.\n\n## The Model Idea and difficulties faced:\nSo what did *I* do? My model idea is quite simple: **8 bagged LGBM models fed into an FTRL model**. Kaggle CPU had memory restrictions of 16GB. Even concatenating all the features together threw that horrible OOM error. I created features piece by piece and concatenated them into three different chunks to create the training data. \n\nNow we have 100M data containing 38 Features. Training an LGB on this large data would still throw an OOM. So I decided to do bagging (which would still throw an OOM error if you didn't slowly convert the data into float numpy arrays yourself using loops). Now Bagging is not an option you'd go for if you wish to make more accurate predictions, but a discussion mentioned that training on larger dataset gave a score boost of ~0.01 and so I did. \n\nTraining on the first chunk (25M) gave a CV of 0.791. After 7 such chunks, I got a CV of 0.792. Frustrated by seeing my LB ranks drop and to improve the predictive power a bit, I decided to add another model to the mix, a FTRL. \n\nI was also motivated to use FTRL for two reasons: \n\n1. I only had two days left before competition ends & FTRL was super fast!\n2. FTRL model is best suited for applications were the predictions need to dynamic and evolving. \n\nFTRL by itself didn't perform very well on my CVs. So I decided to feed my LGB model predictions to it to help it make its predictions. If the FTRL sees the LGB making false predictions during inference, it can improve or alter its predictions accordingly. Doing this improved my LB from 0.792 to 0.793. I have no prior knowledge with FTRL models. Perhaps with better features FTRL might have worked even better.\n\n## Some of my personal takeaways:\n\n- *Don't shy away from newer unfamiliar concepts!* With such a huge time frame given, I could have learned and implemented a transformer model. Instead I chose to stick with what I knew best: Brute forcing my way with hand coded features. Although this approach taught me how to make the most out of the limited given resources.\n- *Numba is awesome!* Ton of code speed up, all with a simple decorator! Numba although fast, may throw OOM if blindly used. My notebook (soon to be published) demonstrates how I leveraged this awesome library.\n- *LB ranks doesn't matter!* To be honest, It was painful to see my LB ranks drop from 60 all the way down to 158. But in the end, I managed to accept it and even laugh about it! See [this](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/208356) hilarious meme thread.\n\n***\nI enjoyed participating in this competition. It gave me so much more than I imagined it would. To all the winners and participants of this competition: *My heartiest congratulations and a Happy new year!*\n***",
    "1143863": "Nice Idea to use FTRLs!",
    "1144381": "My notebook is now public. You can check them out [here](https://www.kaggle.com/doctorkael/riiid-lgb-ftrl-data-gen-and-training-logic)!"
  },
  "source": "meta"
}