{
  "id": 207018,
  "title": "Tips on pipeline engineering for beginners",
  "url": "/competitions/riiid-test-answer-prediction/discussion/207018",
  "author_name": "Jacky",
  "post_date": "2020-12-27T17:27:54.864000",
  "votes": 11,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hey all !</p>\n<p>So I see a lot of people struggeling with the loop feature engineering, I would like to share with you my personnal pipeline that help me avoiding leakages and bugs for submission.</p>\n<p>My pipeline is based on the train/test df as well as two cache files :</p>\n<ul>\n<li>question_cache that is precomputed using the full history and gather information about questions (difficulty, etc…)</li>\n<li>user_cache that is computed iterativaly to avoid any time leakage.</li>\n</ul>\n<p>My pipeline below:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1161354%2F837bee55a510aed5293ba08e8be31d51%2FSans%20titre.png?generation=1609089826464525&amp;alt=media\" alt=\"\"></p>\n<p>In the general sequence, I loop through the row of the train/test set and add the relevant informations from my two caches to build the feature matrix used to train my model. Two things to note:</p>\n<ul>\n<li>It is important to create the feature matrix BEFORE updating the user_cache to avoid leakage of information not yet seen (answered_correctly for example)</li>\n<li>It is very important to consider the questions by batches, as the data is served like this through the API. To make sure there is no batch leaks, I include in my pipeline a temporary file where I store the rows from a unique batch. Only when the batch change, I iterate over this template to first create all the row in the feature matrix and then only update the user_cache.</li>\n<li>Another very important point: Before starting the predictions, you shall precompute the user_cache using the full train df. A lot of information is available from users from the test set in the train set. I was initially not doing it and was stuck in the 0.77X area. This allowed me to jump in the 0.78X one.</li>\n</ul>\n<p>Good luck all for the last few days!</p>",
  "messages": [
    {
      "id": 1128741,
      "postDate": "2020-12-27T17:27:54.863Z",
      "content": "<p>Hey all !</p>\n<p>So I see a lot of people struggeling with the loop feature engineering, I would like to share with you my personnal pipeline that help me avoiding leakages and bugs for submission.</p>\n<p>My pipeline is based on the train/test df as well as two cache files :</p>\n<ul>\n<li>question_cache that is precomputed using the full history and gather information about questions (difficulty, etc…)</li>\n<li>user_cache that is computed iterativaly to avoid any time leakage.</li>\n</ul>\n<p>My pipeline below:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1161354%2F837bee55a510aed5293ba08e8be31d51%2FSans%20titre.png?generation=1609089826464525&amp;alt=media\" alt=\"\"></p>\n<p>In the general sequence, I loop through the row of the train/test set and add the relevant informations from my two caches to build the feature matrix used to train my model. Two things to note:</p>\n<ul>\n<li>It is important to create the feature matrix BEFORE updating the user_cache to avoid leakage of information not yet seen (answered_correctly for example)</li>\n<li>It is very important to consider the questions by batches, as the data is served like this through the API. To make sure there is no batch leaks, I include in my pipeline a temporary file where I store the rows from a unique batch. Only when the batch change, I iterate over this template to first create all the row in the feature matrix and then only update the user_cache.</li>\n<li>Another very important point: Before starting the predictions, you shall precompute the user_cache using the full train df. A lot of information is available from users from the test set in the train set. I was initially not doing it and was stuck in the 0.77X area. This allowed me to jump in the 0.78X one.</li>\n</ul>\n<p>Good luck all for the last few days!</p>",
      "rawMarkdown": "Hey all !\n\nSo I see a lot of people struggeling with the loop feature engineering, I would like to share with you my personnal pipeline that help me avoiding leakages and bugs for submission.\n\nMy pipeline is based on the train/test df as well as two cache files :\n- question_cache that is precomputed using the full history and gather information about questions (difficulty, etc...)\n- user_cache that is computed iterativaly to avoid any time leakage.\n\nMy pipeline below:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1161354%2F837bee55a510aed5293ba08e8be31d51%2FSans%20titre.png?generation=1609089826464525&alt=media)\n\nIn the general sequence, I loop through the row of the train/test set and add the relevant informations from my two caches to build the feature matrix used to train my model. Two things to note:\n- It is important to create the feature matrix BEFORE updating the user_cache to avoid leakage of information not yet seen (answered_correctly for example)\n- It is very important to consider the questions by batches, as the data is served like this through the API. To make sure there is no batch leaks, I include in my pipeline a temporary file where I store the rows from a unique batch. Only when the batch change, I iterate over this template to first create all the row in the feature matrix and then only update the user_cache.\n- Another very important point: Before starting the predictions, you shall precompute the user_cache using the full train df. A lot of information is available from users from the test set in the train set. I was initially not doing it and was stuck in the 0.77X area. This allowed me to jump in the 0.78X one.\n\nGood luck all for the last few days!",
      "votes": 10
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1128741": "Hey all !\n\nSo I see a lot of people struggeling with the loop feature engineering, I would like to share with you my personnal pipeline that help me avoiding leakages and bugs for submission.\n\nMy pipeline is based on the train/test df as well as two cache files :\n- question_cache that is precomputed using the full history and gather information about questions (difficulty, etc...)\n- user_cache that is computed iterativaly to avoid any time leakage.\n\nMy pipeline below:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1161354%2F837bee55a510aed5293ba08e8be31d51%2FSans%20titre.png?generation=1609089826464525&alt=media)\n\nIn the general sequence, I loop through the row of the train/test set and add the relevant informations from my two caches to build the feature matrix used to train my model. Two things to note:\n- It is important to create the feature matrix BEFORE updating the user_cache to avoid leakage of information not yet seen (answered_correctly for example)\n- It is very important to consider the questions by batches, as the data is served like this through the API. To make sure there is no batch leaks, I include in my pipeline a temporary file where I store the rows from a unique batch. Only when the batch change, I iterate over this template to first create all the row in the feature matrix and then only update the user_cache.\n- Another very important point: Before starting the predictions, you shall precompute the user_cache using the full train df. A lot of information is available from users from the test set in the train set. I was initially not doing it and was stuck in the 0.77X area. This allowed me to jump in the 0.78X one.\n\nGood luck all for the last few days!"
  }
}