{
  "id": 201687,
  "title": "Let‘s join Riiids ------ a summary.",
  "url": "/competitions/riiid-test-answer-prediction/discussion/201687",
  "author_name": "biubiuG",
  "post_date": "2020-12-06T06:31:14.862000",
  "votes": 17,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I just enter this competition for two days, some great kernals and discussions are summaried below, I will keep updated, upvote if you like.</p>\n<h1>Riiid</h1>\n<h2>Great EDA Kernals</h2>\n<ol>\n<li><p><a href=\"https://www.kaggle.com/isaienkov/riiid-answer-correctness-prediction-eda-modeling\" target=\"_blank\">Riiid! Answer Correctness Prediction EDA. Modeling</a></p></li>\n<li><p><a href=\"https://www.kaggle.com/erikbruin/riiid-comprehensive-eda-baseline\" target=\"_blank\"><strong>Riiid: Comprehensive EDA + Baseline</strong></a></p>\n<p>this kernal give very impressive information likes below:</p>\n<ul>\n<li><p>every user_id has a content at timestamp = 0</p></li>\n<li><p>“I also want to find out if there is a relationship between timestamp and answered_correctly. To find out I have made 5 bins of timestamp. As you can see, the only noticable thing is that users who have registered relatively recently perform a little worse than users who are active longer.”</p></li></ul></li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5705651%2Ff37c311c86e69fc8d46a99683073cdd5%2F__results___32_0.png?generation=1607236401017999&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>he alse show us the Histogram of percent_correct grouped by task_container_id likes below:</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5705651%2Ff87c9ec60705da4eb333ba8e63bc6317%2F__results___34_0.png?generation=1607236493749848&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>“Below I am plotting the number of answers per user_id against the percentage of questions answered correctly (sample of 200). As some users have answered huge amounts of questions, I have taken out the outliers (user_ids with 1000+ questions answered). As you can see, the trend is upward but there is also a lot of variation among users that have answered few questions.”</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5705651%2F9a47cf1bcf7b63acfe06325be4d7577d%2F__results___37_0.png?generation=1607236520101401&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li><p>the percent answered correctly is about 17% higher when there was an explanation（prior_question_had_explanation = True）. </p></li>\n<li><p>different tags have very different percentofcorrentness, so tags are import features. So does Part.</p></li>\n<li><p>the hidden test set contains new users but not new questions</p></li>\n<li><p>The test data follows chronologically after the train data. The test iterations give interactions of users chronologically.</p></li>\n<li><p>According to <a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">cv-strategy</a>,just taking the last couple of questions from each user as the validation set leads to much on \"light users\"</p>\n<p>See more if you are interested.</p></li>\n</ul>\n<h2>CV stategy</h2>\n<p>cv is important in almost all the competitions. For this one there are many stategies.</p>\n<ol>\n<li><p><a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">Using Last  several entry</a></p>\n<p>Using last several entry for each user as validation data is easy and doesn't look too bad. However, this split method may be focusing too much on light users over heavy users. As a result, the average percentage of correct answers become lower, and there may be a risk of leading us in the wrong direction.</p></li>\n</ol>\n<h2>Model</h2>\n<ol>\n<li><p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/193250\" target=\"_blank\">SAINT+: A Transformer-based model for correctness prediction</a> without code but official</p></li>\n<li><p><a href=\"https://www.kaggle.com/markwijkhuizen/riiid-training-and-prediction-using-a-state\" target=\"_blank\"><strong>Riiid! Training and Prediction using a state</strong></a> LB:0.766（highest score but not for copy and rerun.）</p></li>\n<li><p><a href=\"https://www.kaggle.com/its7171/lgbm-with-loop-feature-engineering\" target=\"_blank\">LGBM with Loop Feature Engineering</a> LB 0.76</p></li>\n<li><p><a href=\"https://www.kaggle.com/erikbruin/riiid-comprehensive-eda-baseline\" target=\"_blank\">Riiid: Comprehensive EDA + Baseline</a> most vote.</p></li>\n</ol>\n<h2>OOM</h2>\n<p>the dataset is so large that may easily cause out of memory(OOM) errors, some discussion talked about it.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/198245\" target=\"_blank\">Reduce memory spike while creating lightgbm dataset</a></li>\n</ul>",
  "messages": [
    {
      "id": 1103701,
      "postDate": "2020-12-06T06:31:14.863Z",
      "content": "<p>I just enter this competition for two days, some great kernals and discussions are summaried below, I will keep updated, upvote if you like.</p>\n<h1>Riiid</h1>\n<h2>Great EDA Kernals</h2>\n<ol>\n<li><p><a href=\"https://www.kaggle.com/isaienkov/riiid-answer-correctness-prediction-eda-modeling\" target=\"_blank\">Riiid! Answer Correctness Prediction EDA. Modeling</a></p></li>\n<li><p><a href=\"https://www.kaggle.com/erikbruin/riiid-comprehensive-eda-baseline\" target=\"_blank\"><strong>Riiid: Comprehensive EDA + Baseline</strong></a></p>\n<p>this kernal give very impressive information likes below:</p>\n<ul>\n<li><p>every user_id has a content at timestamp = 0</p></li>\n<li><p>“I also want to find out if there is a relationship between timestamp and answered_correctly. To find out I have made 5 bins of timestamp. As you can see, the only noticable thing is that users who have registered relatively recently perform a little worse than users who are active longer.”</p></li></ul></li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5705651%2Ff37c311c86e69fc8d46a99683073cdd5%2F__results___32_0.png?generation=1607236401017999&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>he alse show us the Histogram of percent_correct grouped by task_container_id likes below:</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5705651%2Ff87c9ec60705da4eb333ba8e63bc6317%2F__results___34_0.png?generation=1607236493749848&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>“Below I am plotting the number of answers per user_id against the percentage of questions answered correctly (sample of 200). As some users have answered huge amounts of questions, I have taken out the outliers (user_ids with 1000+ questions answered). As you can see, the trend is upward but there is also a lot of variation among users that have answered few questions.”</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5705651%2F9a47cf1bcf7b63acfe06325be4d7577d%2F__results___37_0.png?generation=1607236520101401&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li><p>the percent answered correctly is about 17% higher when there was an explanation（prior_question_had_explanation = True）. </p></li>\n<li><p>different tags have very different percentofcorrentness, so tags are import features. So does Part.</p></li>\n<li><p>the hidden test set contains new users but not new questions</p></li>\n<li><p>The test data follows chronologically after the train data. The test iterations give interactions of users chronologically.</p></li>\n<li><p>According to <a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">cv-strategy</a>,just taking the last couple of questions from each user as the validation set leads to much on \"light users\"</p>\n<p>See more if you are interested.</p></li>\n</ul>\n<h2>CV stategy</h2>\n<p>cv is important in almost all the competitions. For this one there are many stategies.</p>\n<ol>\n<li><p><a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">Using Last  several entry</a></p>\n<p>Using last several entry for each user as validation data is easy and doesn't look too bad. However, this split method may be focusing too much on light users over heavy users. As a result, the average percentage of correct answers become lower, and there may be a risk of leading us in the wrong direction.</p></li>\n</ol>\n<h2>Model</h2>\n<ol>\n<li><p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/193250\" target=\"_blank\">SAINT+: A Transformer-based model for correctness prediction</a> without code but official</p></li>\n<li><p><a href=\"https://www.kaggle.com/markwijkhuizen/riiid-training-and-prediction-using-a-state\" target=\"_blank\"><strong>Riiid! Training and Prediction using a state</strong></a> LB:0.766（highest score but not for copy and rerun.）</p></li>\n<li><p><a href=\"https://www.kaggle.com/its7171/lgbm-with-loop-feature-engineering\" target=\"_blank\">LGBM with Loop Feature Engineering</a> LB 0.76</p></li>\n<li><p><a href=\"https://www.kaggle.com/erikbruin/riiid-comprehensive-eda-baseline\" target=\"_blank\">Riiid: Comprehensive EDA + Baseline</a> most vote.</p></li>\n</ol>\n<h2>OOM</h2>\n<p>the dataset is so large that may easily cause out of memory(OOM) errors, some discussion talked about it.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/198245\" target=\"_blank\">Reduce memory spike while creating lightgbm dataset</a></li>\n</ul>",
      "rawMarkdown": "I just enter this competition for two days, some great kernals and discussions are summaried below, I will keep updated, upvote if you like.\n# Riiid\n##Great EDA Kernals\n\n1. [Riiid! Answer Correctness Prediction EDA. Modeling](https://www.kaggle.com/isaienkov/riiid-answer-correctness-prediction-eda-modeling)\n\n2. [**Riiid: Comprehensive EDA + Baseline**](https://www.kaggle.com/erikbruin/riiid-comprehensive-eda-baseline)\n\n   this kernal give very impressive information likes below:\n\n   - every user_id has a content at timestamp = 0\n\n   - “I also want to find out if there is a relationship between timestamp and answered_correctly. To find out I have made 5 bins of timestamp. As you can see, the only noticable thing is that users who have registered relatively recently perform a little worse than users who are active longer.”\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5705651%2Ff37c311c86e69fc8d46a99683073cdd5%2F__results___32_0.png?generation=1607236401017999&alt=media)\n\n   - he alse show us the Histogram of percent_correct grouped by task_container_id likes below:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5705651%2Ff87c9ec60705da4eb333ba8e63bc6317%2F__results___34_0.png?generation=1607236493749848&alt=media)\n\n   - “Below I am plotting the number of answers per user_id against the percentage of questions answered correctly (sample of 200). As some users have answered huge amounts of questions, I have taken out the outliers (user_ids with 1000+ questions answered). As you can see, the trend is upward but there is also a lot of variation among users that have answered few questions.”\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5705651%2F9a47cf1bcf7b63acfe06325be4d7577d%2F__results___37_0.png?generation=1607236520101401&alt=media)\n\n   - the percent answered correctly is about 17% higher when there was an explanation（prior_question_had_explanation = True）. \n\n   - different tags have very different percentofcorrentness, so tags are import features. So does Part.\n\n   - the hidden test set contains new users but not new questions\n\n   - The test data follows chronologically after the train data. The test iterations give interactions of users chronologically.\n\n   - According to [cv-strategy](https://www.kaggle.com/its7171/cv-strategy),just taking the last couple of questions from each user as the validation set leads to much on \"light users\"\n\n   See more if you are interested.\n\n## CV stategy\n\n   cv is important in almost all the competitions. For this one there are many stategies.\n\n   1. [Using Last  several entry](https://www.kaggle.com/its7171/cv-strategy)\n\n   Using last several entry for each user as validation data is easy and doesn't look too bad. However, this split method may be focusing too much on light users over heavy users. As a result, the average percentage of correct answers become lower, and there may be a risk of leading us in the wrong direction.\n\n## Model\n\n   1. [SAINT+: A Transformer-based model for correctness prediction](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/193250) without code but official\n\n   2. [**Riiid! Training and Prediction using a state**](https://www.kaggle.com/markwijkhuizen/riiid-training-and-prediction-using-a-state) LB:0.766（highest score but not for copy and rerun.）\n\n   3. [LGBM with Loop Feature Engineering](https://www.kaggle.com/its7171/lgbm-with-loop-feature-engineering) LB 0.76\n\n   4. [Riiid: Comprehensive EDA + Baseline](https://www.kaggle.com/erikbruin/riiid-comprehensive-eda-baseline) most vote.\n## OOM\nthe dataset is so large that may easily cause out of memory(OOM) errors, some discussion talked about it.\n- [Reduce memory spike while creating lightgbm dataset](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/198245)",
      "votes": 17
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1103701": "I just enter this competition for two days, some great kernals and discussions are summaried below, I will keep updated, upvote if you like.\n# Riiid\n##Great EDA Kernals\n\n1. [Riiid! Answer Correctness Prediction EDA. Modeling](https://www.kaggle.com/isaienkov/riiid-answer-correctness-prediction-eda-modeling)\n\n2. [**Riiid: Comprehensive EDA + Baseline**](https://www.kaggle.com/erikbruin/riiid-comprehensive-eda-baseline)\n\n   this kernal give very impressive information likes below:\n\n   - every user_id has a content at timestamp = 0\n\n   - “I also want to find out if there is a relationship between timestamp and answered_correctly. To find out I have made 5 bins of timestamp. As you can see, the only noticable thing is that users who have registered relatively recently perform a little worse than users who are active longer.”\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5705651%2Ff37c311c86e69fc8d46a99683073cdd5%2F__results___32_0.png?generation=1607236401017999&alt=media)\n\n   - he alse show us the Histogram of percent_correct grouped by task_container_id likes below:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5705651%2Ff87c9ec60705da4eb333ba8e63bc6317%2F__results___34_0.png?generation=1607236493749848&alt=media)\n\n   - “Below I am plotting the number of answers per user_id against the percentage of questions answered correctly (sample of 200). As some users have answered huge amounts of questions, I have taken out the outliers (user_ids with 1000+ questions answered). As you can see, the trend is upward but there is also a lot of variation among users that have answered few questions.”\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5705651%2F9a47cf1bcf7b63acfe06325be4d7577d%2F__results___37_0.png?generation=1607236520101401&alt=media)\n\n   - the percent answered correctly is about 17% higher when there was an explanation（prior_question_had_explanation = True）. \n\n   - different tags have very different percentofcorrentness, so tags are import features. So does Part.\n\n   - the hidden test set contains new users but not new questions\n\n   - The test data follows chronologically after the train data. The test iterations give interactions of users chronologically.\n\n   - According to [cv-strategy](https://www.kaggle.com/its7171/cv-strategy),just taking the last couple of questions from each user as the validation set leads to much on \"light users\"\n\n   See more if you are interested.\n\n## CV stategy\n\n   cv is important in almost all the competitions. For this one there are many stategies.\n\n   1. [Using Last  several entry](https://www.kaggle.com/its7171/cv-strategy)\n\n   Using last several entry for each user as validation data is easy and doesn't look too bad. However, this split method may be focusing too much on light users over heavy users. As a result, the average percentage of correct answers become lower, and there may be a risk of leading us in the wrong direction.\n\n## Model\n\n   1. [SAINT+: A Transformer-based model for correctness prediction](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/193250) without code but official\n\n   2. [**Riiid! Training and Prediction using a state**](https://www.kaggle.com/markwijkhuizen/riiid-training-and-prediction-using-a-state) LB:0.766（highest score but not for copy and rerun.）\n\n   3. [LGBM with Loop Feature Engineering](https://www.kaggle.com/its7171/lgbm-with-loop-feature-engineering) LB 0.76\n\n   4. [Riiid: Comprehensive EDA + Baseline](https://www.kaggle.com/erikbruin/riiid-comprehensive-eda-baseline) most vote.\n## OOM\nthe dataset is so large that may easily cause out of memory(OOM) errors, some discussion talked about it.\n- [Reduce memory spike while creating lightgbm dataset](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/198245)"
  }
}