{
  "id": 209816,
  "title": "My personnal journey through the RIID competition - Silver LGB single model",
  "url": "/competitions/riiid-test-answer-prediction/writeups/jacky-my-personnal-journey-through-the-riid-compet",
  "author_name": "",
  "post_date": "2021-01-08T17:46:34.567Z",
  "votes": 13,
  "comment_count": 3,
  "views": 0,
  "content": "<p>First of all, and as many others, I had a lot of pleasure participating to that competition. The challenge was fun and there were many challenges that I was not used to face. </p>\n<p>I joined the competition a bit late (1 month before the end) so I really decided to focus on the feature engineering part. I had the aim of reaching the silver zone, and that's what I did, so very happy about it !</p>\n<p>Overall I learned a lot of tricks to handle RAM and improve CPU computational time:</p>\n<ul>\n<li>Making a single loop to make the whole feature engineering rather than using several pandas.apply()</li>\n<li>Using a csv iterator rather than iterating through a numpy array fully loaded (or worst, a dataframe!)</li>\n<li>Smart use of sets to calculate presence or not of elements in a list</li>\n<li>The potential of dictionaries that I clearly underestimated until now !</li>\n<li>Some RAM tricks such than converting the training array column by column to float32 before feeding it to the LGB model and avoid RAM explosion</li>\n<li>The potential of SHAP for feature selection</li>\n<li>The possibility to use conditional features rather than doing One hot Encoding</li>\n<li>…</li>\n</ul>\n<p>For me, one of the key challenges was to not get lost in the amount of code I was producing, and it happened more than once that I decided to start again from a blank page to have a clear view of what I was actually doing ! If there is something I will try to handle better next time it is really global code organization. I think I lost a lot (too much) on it.</p>\n<p>Separating the code in 3 parts worked very well for me, with one part handling creation of user history, one part updating it, and one part creating the features based on the user history.</p>\n<p>Apart from all classical features that have been already discussed, one of my model key differentiation was the bayesian probability index that I built and discussed <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/207148\" target=\"_blank\">there</a>. Following <a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> comment, I created an efficient cluster-based feature, illustrated by the feature below:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1161354%2F9a1323fb5376e39469c104cae83ac379%2Fclusterriid.png?generation=1610124766536795&amp;alt=media\" alt=\"\"></p>\n<p>For those interested, I created a detailed kernel with all my inference pipeline and feature engineering. You can check it out here, there is probably one or two thing interesting to pick up !<br>\n<a href=\"https://www.kaggle.com/bowaka/riid-lgb-single-model-0-793-full-summary\" target=\"_blank\">Notebook here</a></p>\n<p>I will personally continuing to work on the datasets and take the opportunity to improve my knowledge and expertise about transformers (that is currently close to 0 !)</p>",
  "messages": [
    {
      "id": "1144782",
      "postDate": "01/08/2021 17:07:05",
      "content": "<p>First of all, and as many others, I had a lot of pleasure participating to that competition. The challenge was fun and there were many challenges that I was not used to face. </p>\n<p>I joined the competition a bit late (1 month before the end) so I really decided to focus on the feature engineering part. I had the aim of reaching the silver zone, and that's what I did, so very happy about it !</p>\n<p>Overall I learned a lot of tricks to handle RAM and improve CPU computational time:</p>\n<ul>\n<li>Making a single loop to make the whole feature engineering rather than using several pandas.apply()</li>\n<li>Using a csv iterator rather than iterating through a numpy array fully loaded (or worst, a dataframe!)</li>\n<li>Smart use of sets to calculate presence or not of elements in a list</li>\n<li>The potential of dictionaries that I clearly underestimated until now !</li>\n<li>Some RAM tricks such than converting the training array column by column to float32 before feeding it to the LGB model and avoid RAM explosion</li>\n<li>The potential of SHAP for feature selection</li>\n<li>The possibility to use conditional features rather than doing One hot Encoding</li>\n<li>…</li>\n</ul>\n<p>For me, one of the key challenges was to not get lost in the amount of code I was producing, and it happened more than once that I decided to start again from a blank page to have a clear view of what I was actually doing ! If there is something I will try to handle better next time it is really global code organization. I think I lost a lot (too much) on it.</p>\n<p>Separating the code in 3 parts worked very well for me, with one part handling creation of user history, one part updating it, and one part creating the features based on the user history.</p>\n<p>Apart from all classical features that have been already discussed, one of my model key differentiation was the bayesian probability index that I built and discussed <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/207148\" target=\"_blank\">there</a>. Following <a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> comment, I created an efficient cluster-based feature, illustrated by the feature below:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1161354%2F9a1323fb5376e39469c104cae83ac379%2Fclusterriid.png?generation=1610124766536795&amp;alt=media\" alt=\"\"></p>\n<p>For those interested, I created a detailed kernel with all my inference pipeline and feature engineering. You can check it out here, there is probably one or two thing interesting to pick up !<br>\n<a href=\"https://www.kaggle.com/bowaka/riid-lgb-single-model-0-793-full-summary\" target=\"_blank\">Notebook here</a></p>\n<p>I will personally continuing to work on the datasets and take the opportunity to improve my knowledge and expertise about transformers (that is currently close to 0 !)</p>",
      "rawMarkdown": "First of all, and as many others, I had a lot of pleasure participating to that competition. The challenge was fun and there were many challenges that I was not used to face. \n\nI joined the competition a bit late (1 month before the end) so I really decided to focus on the feature engineering part. I had the aim of reaching the silver zone, and that's what I did, so very happy about it !\n\nOverall I learned a lot of tricks to handle RAM and improve CPU computational time:\n- Making a single loop to make the whole feature engineering rather than using several pandas.apply()\n- Using a csv iterator rather than iterating through a numpy array fully loaded (or worst, a dataframe!)\n- Smart use of sets to calculate presence or not of elements in a list\n- The potential of dictionaries that I clearly underestimated until now !\n- Some RAM tricks such than converting the training array column by column to float32 before feeding it to the LGB model and avoid RAM explosion\n- The potential of SHAP for feature selection\n- The possibility to use conditional features rather than doing One hot Encoding\n- ...\n\nFor me, one of the key challenges was to not get lost in the amount of code I was producing, and it happened more than once that I decided to start again from a blank page to have a clear view of what I was actually doing ! If there is something I will try to handle better next time it is really global code organization. I think I lost a lot (too much) on it.\n\nSeparating the code in 3 parts worked very well for me, with one part handling creation of user history, one part updating it, and one part creating the features based on the user history.\n\nApart from all classical features that have been already discussed, one of my model key differentiation was the bayesian probability index that I built and discussed [there](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/207148). Following @authman comment, I created an efficient cluster-based feature, illustrated by the feature below:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1161354%2F9a1323fb5376e39469c104cae83ac379%2Fclusterriid.png?generation=1610124766536795&alt=media)\n\nFor those interested, I created a detailed kernel with all my inference pipeline and feature engineering. You can check it out here, there is probably one or two thing interesting to pick up !\n[Notebook here](https://www.kaggle.com/bowaka/riid-lgb-single-model-0-793-full-summary)\n\nI will personally continuing to work on the datasets and take the opportunity to improve my knowledge and expertise about transformers (that is currently close to 0 !)",
      "votes": null
    },
    {
      "id": "1144956",
      "postDate": "01/08/2021 19:17:19",
      "content": "<p>congratulations, my lgb 0.7928 maybe, two rank to silver …transformer really work in this competition, my saint 0.78, with 3 features.though I can't ensemble with my lgb model because the RAM.I will continue my saint model in this dataset too.</p>",
      "rawMarkdown": "congratulations, my lgb 0.7928 maybe, two rank to silver …transformer really work in this competition, my saint 0.78, with 3 features.though I can't ensemble with my lgb model because the RAM.I will continue my saint model in this dataset too.",
      "votes": null
    },
    {
      "id": "1144965",
      "postDate": "01/08/2021 19:21:21",
      "content": "<p>I will read your notebook tomorrow, learn handle RAM . this is my lacking skill</p>",
      "rawMarkdown": "I will read your notebook tomorrow, learn handle RAM . this is my lacking skill",
      "votes": null
    },
    {
      "id": "1145039",
      "postDate": "01/08/2021 20:28:40",
      "content": "<p>Yes I saw you in the leaderboard, I'm sorry you missed the silver from so few… Next will be the good one! </p>",
      "rawMarkdown": "Yes I saw you in the leaderboard, I'm sorry you missed the silver from so few... Next will be the good one!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1144956,
      "author_name": "yangxiaoshuai",
      "author_url": "",
      "post_date": "01/08/2021 19:17:19",
      "content": "<p>congratulations, my lgb 0.7928 maybe, two rank to silver …transformer really work in this competition, my saint 0.78, with 3 features.though I can't ensemble with my lgb model because the RAM.I will continue my saint model in this dataset too.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1145039,
          "author_name": "bowaka",
          "author_url": "",
          "post_date": "01/08/2021 20:28:40",
          "content": "<p>Yes I saw you in the leaderboard, I'm sorry you missed the silver from so few… Next will be the good one! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1144965,
      "author_name": "yangxiaoshuai",
      "author_url": "",
      "post_date": "01/08/2021 19:21:21",
      "content": "<p>I will read your notebook tomorrow, learn handle RAM . this is my lacking skill</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1144782": "First of all, and as many others, I had a lot of pleasure participating to that competition. The challenge was fun and there were many challenges that I was not used to face. \n\nI joined the competition a bit late (1 month before the end) so I really decided to focus on the feature engineering part. I had the aim of reaching the silver zone, and that's what I did, so very happy about it !\n\nOverall I learned a lot of tricks to handle RAM and improve CPU computational time:\n- Making a single loop to make the whole feature engineering rather than using several pandas.apply()\n- Using a csv iterator rather than iterating through a numpy array fully loaded (or worst, a dataframe!)\n- Smart use of sets to calculate presence or not of elements in a list\n- The potential of dictionaries that I clearly underestimated until now !\n- Some RAM tricks such than converting the training array column by column to float32 before feeding it to the LGB model and avoid RAM explosion\n- The potential of SHAP for feature selection\n- The possibility to use conditional features rather than doing One hot Encoding\n- ...\n\nFor me, one of the key challenges was to not get lost in the amount of code I was producing, and it happened more than once that I decided to start again from a blank page to have a clear view of what I was actually doing ! If there is something I will try to handle better next time it is really global code organization. I think I lost a lot (too much) on it.\n\nSeparating the code in 3 parts worked very well for me, with one part handling creation of user history, one part updating it, and one part creating the features based on the user history.\n\nApart from all classical features that have been already discussed, one of my model key differentiation was the bayesian probability index that I built and discussed [there](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/207148). Following @authman comment, I created an efficient cluster-based feature, illustrated by the feature below:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1161354%2F9a1323fb5376e39469c104cae83ac379%2Fclusterriid.png?generation=1610124766536795&alt=media)\n\nFor those interested, I created a detailed kernel with all my inference pipeline and feature engineering. You can check it out here, there is probably one or two thing interesting to pick up !\n[Notebook here](https://www.kaggle.com/bowaka/riid-lgb-single-model-0-793-full-summary)\n\nI will personally continuing to work on the datasets and take the opportunity to improve my knowledge and expertise about transformers (that is currently close to 0 !)",
    "1144956": "congratulations, my lgb 0.7928 maybe, two rank to silver …transformer really work in this competition, my saint 0.78, with 3 features.though I can't ensemble with my lgb model because the RAM.I will continue my saint model in this dataset too.",
    "1144965": "I will read your notebook tomorrow, learn handle RAM . this is my lacking skill",
    "1145039": "Yes I saw you in the leaderboard, I'm sorry you missed the silver from so few... Next will be the good one!"
  },
  "source": "meta"
}