{
  "id": 209474,
  "title": "Why my LGB score is struck at 0.785 ? ",
  "url": "/competitions/riiid-test-answer-prediction/discussion/209474",
  "author_name": "",
  "post_date": "2021-01-07T16:57:44.425761800Z",
  "votes": 2,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Hi, <br>\n   I used the following features for the lgb model, Total (49 features )</p>\n<p><strong>continuous variables :</strong></p>\n<p>user_correct_mean</p>\n<p>question_mean</p>\n<p>tags_encoded_mean - questions having the same set of tags are grouped together and the correct mean for them. </p>\n<p>lecture_tag_seen - number of lectures having tag which is in the question / number of tags in ques.</p>\n<p>question_tag_seen - sum of number of times explanation seen for the question having tag belonging to the one in the current question / number of tags of the question</p>\n<p>user_question_count</p>\n<p>question_count</p>\n<p>tags_encoded_count</p>\n<p>'timestamp_recency_1'<br>\n'timestamp_recency_2'<br>\n'timestamp_recency_3', </p>\n<p>'user_prior_question_had_explanation_mean' - mean of number of times explanation seen by user</p>\n<p>'user_prior_question_elapsed_time_mean' - mean time the user takes to. answer question</p>\n<p>prior_question_elapsed_time</p>\n<p>number_of_watched_lectures</p>\n<p>average_time_between_interactions - timestamp / number of interactions</p>\n<p>last_correct_recency</p>\n<p>last_incorrect_recency</p>\n<p>continuous_incorrect</p>\n<p>continuous_correct</p>\n<p>hmean_content_id</p>\n<p>hmean_tags_encoded_id</p>\n<p>answered_correctly_rollsum - number of correct_answers for last 10 task container ids</p>\n<p>answered_mean_rollsum  - similar to above but mean</p>\n<p>repeated_content_count </p>\n<p>repeated_content_mean</p>\n<p>repeated_part_count </p>\n<p>repeated_part_mean</p>\n<p>bundle_mean</p>\n<p>bundle_elapsed_time_mean - avg time taken by all users to answer this bundle</p>\n<p>bundle_explanation_mean - avg no of times explanation seen by user for this bundle(btw 0 to 1)</p>\n<p>max_difficulty_crt</p>\n<p>min_difficulty_wrng</p>\n<p>user_correct_std</p>\n<p>part_max_difficulty_crt</p>\n<p>part_min_difficulty_wrng</p>\n<p>number_of_tags</p>\n<p>user_tag_mean_accuracy - avg of mean correctness for. each tag in in. the question</p>\n<p>user_tag_max_accuracy</p>\n<p>user_tag_min_accuracy</p>\n<p>user_rating_mean<br>\nuser_rating_std<br>\nquestion_rating_mean<br>\nquestion_rating_std<br>\ndraw_prob - All these belong to true_skill</p>\n<p>lecture_recency - time interval between the current question and the last seen lecture (1e12 if not seen)</p>\n<p>question_tag_lecture_max_recency - time interval between the each tag of the current question and the last seen lecture with the same tag. Then take the min of those (1e12 if not seen)</p>\n<p>question_tag_lecture_min_recency - similar to above but taking max. </p>\n<p><strong>Categorical:</strong></p>\n<p>question_part</p>\n<p>I updated only after completing a bundle. </p>\n<p>I have seen in the discussion people with fewer features than mine have reached 0.79+. But I am getting only. 0.784 with these features though I dont think none of these are noise. </p>\n<p>I got validation score of 0.785 with this. I have submitted two submissions prior and I have noted that the lb score is atleast 0.005 more than valid score. (I use different validation set - custom i.e Taking last 5% of data for each users as test and plus 2.5% of the whole users as new)</p>\n<p>My last submission with features upto repeated_content_mean ( top 26 features mentioned above ) scored about 0.781. But I dont know where did I go wrong. Any help would be highly appreciated. </p>\n<p>Although I dont expect reply before this competition ends, but knowing the reason will help me a lot moving forward. Thanks. </p>",
  "messages": [
    {
      "id": "1142907",
      "postDate": "01/07/2021 16:57:44",
      "content": "<p>Hi, <br>\n   I used the following features for the lgb model, Total (49 features )</p>\n<p><strong>continuous variables :</strong></p>\n<p>user_correct_mean</p>\n<p>question_mean</p>\n<p>tags_encoded_mean - questions having the same set of tags are grouped together and the correct mean for them. </p>\n<p>lecture_tag_seen - number of lectures having tag which is in the question / number of tags in ques.</p>\n<p>question_tag_seen - sum of number of times explanation seen for the question having tag belonging to the one in the current question / number of tags of the question</p>\n<p>user_question_count</p>\n<p>question_count</p>\n<p>tags_encoded_count</p>\n<p>'timestamp_recency_1'<br>\n'timestamp_recency_2'<br>\n'timestamp_recency_3', </p>\n<p>'user_prior_question_had_explanation_mean' - mean of number of times explanation seen by user</p>\n<p>'user_prior_question_elapsed_time_mean' - mean time the user takes to. answer question</p>\n<p>prior_question_elapsed_time</p>\n<p>number_of_watched_lectures</p>\n<p>average_time_between_interactions - timestamp / number of interactions</p>\n<p>last_correct_recency</p>\n<p>last_incorrect_recency</p>\n<p>continuous_incorrect</p>\n<p>continuous_correct</p>\n<p>hmean_content_id</p>\n<p>hmean_tags_encoded_id</p>\n<p>answered_correctly_rollsum - number of correct_answers for last 10 task container ids</p>\n<p>answered_mean_rollsum  - similar to above but mean</p>\n<p>repeated_content_count </p>\n<p>repeated_content_mean</p>\n<p>repeated_part_count </p>\n<p>repeated_part_mean</p>\n<p>bundle_mean</p>\n<p>bundle_elapsed_time_mean - avg time taken by all users to answer this bundle</p>\n<p>bundle_explanation_mean - avg no of times explanation seen by user for this bundle(btw 0 to 1)</p>\n<p>max_difficulty_crt</p>\n<p>min_difficulty_wrng</p>\n<p>user_correct_std</p>\n<p>part_max_difficulty_crt</p>\n<p>part_min_difficulty_wrng</p>\n<p>number_of_tags</p>\n<p>user_tag_mean_accuracy - avg of mean correctness for. each tag in in. the question</p>\n<p>user_tag_max_accuracy</p>\n<p>user_tag_min_accuracy</p>\n<p>user_rating_mean<br>\nuser_rating_std<br>\nquestion_rating_mean<br>\nquestion_rating_std<br>\ndraw_prob - All these belong to true_skill</p>\n<p>lecture_recency - time interval between the current question and the last seen lecture (1e12 if not seen)</p>\n<p>question_tag_lecture_max_recency - time interval between the each tag of the current question and the last seen lecture with the same tag. Then take the min of those (1e12 if not seen)</p>\n<p>question_tag_lecture_min_recency - similar to above but taking max. </p>\n<p><strong>Categorical:</strong></p>\n<p>question_part</p>\n<p>I updated only after completing a bundle. </p>\n<p>I have seen in the discussion people with fewer features than mine have reached 0.79+. But I am getting only. 0.784 with these features though I dont think none of these are noise. </p>\n<p>I got validation score of 0.785 with this. I have submitted two submissions prior and I have noted that the lb score is atleast 0.005 more than valid score. (I use different validation set - custom i.e Taking last 5% of data for each users as test and plus 2.5% of the whole users as new)</p>\n<p>My last submission with features upto repeated_content_mean ( top 26 features mentioned above ) scored about 0.781. But I dont know where did I go wrong. Any help would be highly appreciated. </p>\n<p>Although I dont expect reply before this competition ends, but knowing the reason will help me a lot moving forward. Thanks. </p>",
      "rawMarkdown": "Hi, \n   I used the following features for the lgb model, Total (49 features )\n\n**continuous variables :**\n\nuser_correct_mean\n\nquestion_mean\n\ntags_encoded_mean - questions having the same set of tags are grouped together and the correct mean for them. \n\nlecture_tag_seen - number of lectures having tag which is in the question / number of tags in ques.\n\nquestion_tag_seen - sum of number of times explanation seen for the question having tag belonging to the one in the current question / number of tags of the question\n\nuser_question_count\n\nquestion_count\n\ntags_encoded_count\n\n'timestamp_recency_1'\n'timestamp_recency_2'\n'timestamp_recency_3', \n\n'user_prior_question_had_explanation_mean' - mean of number of times explanation seen by user\n\n'user_prior_question_elapsed_time_mean' - mean time the user takes to. answer question\n\nprior_question_elapsed_time\n\nnumber_of_watched_lectures\n\naverage_time_between_interactions - timestamp / number of interactions\n\nlast_correct_recency\n\nlast_incorrect_recency\n\ncontinuous_incorrect\n\ncontinuous_correct\n\nhmean_content_id\n\nhmean_tags_encoded_id\n\nanswered_correctly_rollsum - number of correct_answers for last 10 task container ids\n\nanswered_mean_rollsum  - similar to above but mean\n\nrepeated_content_count \n\nrepeated_content_mean\n\nrepeated_part_count \n\nrepeated_part_mean\n\nbundle_mean\n\nbundle_elapsed_time_mean - avg time taken by all users to answer this bundle\n\nbundle_explanation_mean - avg no of times explanation seen by user for this bundle(btw 0 to 1)\n\nmax_difficulty_crt\n\nmin_difficulty_wrng\n\nuser_correct_std\n\npart_max_difficulty_crt\n\npart_min_difficulty_wrng\n\nnumber_of_tags\n\nuser_tag_mean_accuracy - avg of mean correctness for. each tag in in. the question\n\nuser_tag_max_accuracy\n\nuser_tag_min_accuracy\n\nuser_rating_mean\nuser_rating_std\nquestion_rating_mean\nquestion_rating_std\ndraw_prob - All these belong to true_skill\n\nlecture_recency - time interval between the current question and the last seen lecture (1e12 if not seen)\n\nquestion_tag_lecture_max_recency - time interval between the each tag of the current question and the last seen lecture with the same tag. Then take the min of those (1e12 if not seen)\n\nquestion_tag_lecture_min_recency - similar to above but taking max. \n\n\n**Categorical:**\n\nquestion_part\n\nI updated only after completing a bundle. \n\nI have seen in the discussion people with fewer features than mine have reached 0.79+. But I am getting only. 0.784 with these features though I dont think none of these are noise. \n\nI got validation score of 0.785 with this. I have submitted two submissions prior and I have noted that the lb score is atleast 0.005 more than valid score. (I use different validation set - custom i.e Taking last 5% of data for each users as test and plus 2.5% of the whole users as new)\n\nMy last submission with features upto repeated_content_mean ( top 26 features mentioned above ) scored about 0.781. But I dont know where did I go wrong. Any help would be highly appreciated. \n\nAlthough I dont expect reply before this competition ends, but knowing the reason will help me a lot moving forward. Thanks.",
      "votes": null
    },
    {
      "id": "1142955",
      "postDate": "01/07/2021 17:27:23",
      "content": "<p>Hi ! I will personnaly make a detailed presentation of my current LGB that managed to reach 0.792 LB when the competition is over ! Hope it can help 🙂</p>",
      "rawMarkdown": "Hi ! I will personnaly make a detailed presentation of my current LGB that managed to reach 0.792 LB when the competition is over ! Hope it can help 🙂",
      "votes": null
    },
    {
      "id": "1143020",
      "postDate": "01/07/2021 18:08:05",
      "content": "<p>I have an lgbm model with 792 in public leaderboard and 797 in local validation. It will be very interesting to read the write-up of lgbm solutions with 0.8+(if that is possible). How far a boosting model can go with this data? What are the\"magic\" features? Any especial treatment in the data?</p>",
      "rawMarkdown": "I have an lgbm model with 792 in public leaderboard and 797 in local validation. It will be very interesting to read the write-up of lgbm solutions with 0.8+(if that is possible). How far a boosting model can go with this data? What are the\"magic\" features? Any especial treatment in the data?",
      "votes": null
    },
    {
      "id": "1143023",
      "postDate": "01/07/2021 18:09:59",
      "content": "<p>I get same validation and same public ! Curious to see the different write-up too !</p>",
      "rawMarkdown": "I get same validation and same public ! Curious to see the different write-up too !",
      "votes": null
    },
    {
      "id": "1143062",
      "postDate": "01/07/2021 18:32:04",
      "content": "<p>Yeah. Will be very helpful. Thanks<br>\nAlso if you. ensemble with publicly available Transformer model, you will definitely get a boost in score. My 0.005 boost is just  with ensemble of transformer model, publicly available. This one <br>\n<a href=\"https://www.kaggle.com/gilfernandes/riiid-self-attention-transformer\" target=\"_blank\">https://www.kaggle.com/gilfernandes/riiid-self-attention-transformer</a></p>",
      "rawMarkdown": "Yeah. Will be very helpful. Thanks\nAlso if you. ensemble with publicly available Transformer model, you will definitely get a boost in score. My 0.005 boost is just  with ensemble of transformer model, publicly available. This one \nhttps://www.kaggle.com/gilfernandes/riiid-self-attention-transformer",
      "votes": null
    },
    {
      "id": "1143090",
      "postDate": "01/07/2021 18:44:58",
      "content": "<p>my lgb CV 0.795 LB may be 0.792or0.791, there are about five or four important features I think, and more train data score better.</p>",
      "rawMarkdown": "my lgb CV 0.795 LB may be 0.792or0.791, there are about five or four important features I think, and more train data score better.",
      "votes": null
    },
    {
      "id": "1143096",
      "postDate": "01/07/2021 18:51:54",
      "content": "<p>I will be curious to see a model using the top magic features of everybody !</p>",
      "rawMarkdown": "I will be curious to see a model using the top magic features of everybody !",
      "votes": null
    },
    {
      "id": "1143129",
      "postDate": "01/07/2021 19:05:52",
      "content": "<p>Can confirm a single LGB model trained on 1m rows can reach 0.800 LB. If trained on the full dataset I think it could be competitive for a prize by itself. Unfortunately I didn't have the compute to train on more data in time as I had bugs until this morning and it takes a long time to generate the features. But I'm satisfied with the design even if I couldn't push out the best version in the end. </p>\n<p>Looking forward to reading solutions soon. For me, the \"magic\" was to map the competition data to the open source Ednet data, which has some extra information. It added ~0.01. </p>",
      "rawMarkdown": "Can confirm a single LGB model trained on 1m rows can reach 0.800 LB. If trained on the full dataset I think it could be competitive for a prize by itself. Unfortunately I didn't have the compute to train on more data in time as I had bugs until this morning and it takes a long time to generate the features. But I'm satisfied with the design even if I couldn't push out the best version in the end. \n\nLooking forward to reading solutions soon. For me, the \"magic\" was to map the competition data to the open source Ednet data, which has some extra information. It added ~0.01.",
      "votes": null
    },
    {
      "id": "1143186",
      "postDate": "01/07/2021 19:32:20",
      "content": "<p>I can confirm this also and for the same reasons + holiday plans I just couldn't find the time to execute my models properly. The engineering alone is a big part of this competition.</p>",
      "rawMarkdown": "I can confirm this also and for the same reasons + holiday plans I just couldn't find the time to execute my models properly. The engineering alone is a big part of this competition.",
      "votes": null
    },
    {
      "id": "1143438",
      "postDate": "01/07/2021 22:46:56",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5326475%2Fa1a714ace0c8b191c66b3b89b144a3ee%2FFotor_161005881525124.jpg?generation=1610058848758943&amp;alt=media\" alt=\"\"></p>\n<p>Above is my used features, cv about 0.7926.., lb 0.792<br>\nI only use kaggle notebook, using last 29M, sample 15M to train</p>\n<p>upm = user_part_mean<br>\ntb = user_last_samebundle_recency1<br>\ntp = user_last_samepart_recency1<br>\nuccm = user_concent_correctness_mean<br>\ncucm = content_user_correctness_mean<br>\nubcm = user_bundle_correctness_mean<br>\nus = user_correct<br>\nups = user_part_correct</p>\n<p>I am looking forward to see someone's solutions which can achieve 0.795+</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5326475%2Fa1a714ace0c8b191c66b3b89b144a3ee%2FFotor_161005881525124.jpg?generation=1610058848758943&alt=media)\n\nAbove is my used features, cv about 0.7926.., lb 0.792\nI only use kaggle notebook, using last 29M, sample 15M to train\n\nupm = user_part_mean\ntb = user_last_samebundle_recency1\ntp = user_last_samepart_recency1\nuccm = user_concent_correctness_mean\ncucm = content_user_correctness_mean\nubcm = user_bundle_correctness_mean\nus = user_correct\nups = user_part_correct\n\nI am looking forward to see someone's solutions which can achieve 0.795+",
      "votes": null
    },
    {
      "id": "1143447",
      "postDate": "01/07/2021 22:54:09",
      "content": "<p>You used the data (Ednet) other than the training data to achieve 0.800?</p>",
      "rawMarkdown": "You used the data (Ednet) other than the training data to achieve 0.800?",
      "votes": null
    },
    {
      "id": "1143472",
      "postDate": "01/07/2021 23:21:11",
      "content": "<p>I have 0.792 with some another features. Maybe, if add them to your solution it will be 0.795+</p>",
      "rawMarkdown": "I have 0.792 with some another features. Maybe, if add them to your solution it will be 0.795+",
      "votes": null
    },
    {
      "id": "1143484",
      "postDate": "01/07/2021 23:30:18",
      "content": "<p>Yes, the data you can find at <a href=\"https://github.com/riiid/ednet\" target=\"_blank\">https://github.com/riiid/ednet</a></p>",
      "rawMarkdown": "Yes, the data you can find at https://github.com/riiid/ednet",
      "votes": null
    },
    {
      "id": "1143611",
      "postDate": "01/08/2021 01:11:08",
      "content": "<p>I found three magic features at least by top LGBM solutions, all amazing,</p>",
      "rawMarkdown": "I found three magic features at least by top LGBM solutions, all amazing,",
      "votes": null
    },
    {
      "id": "1143944",
      "postDate": "01/08/2021 06:30:39",
      "content": "<p>From my experience, timestamp_recency_4 and timestamp_recency_5 still have some predictive power and improved the score. I add them as long as they improve my model.</p>",
      "rawMarkdown": "From my experience, timestamp_recency_4 and timestamp_recency_5 still have some predictive power and improved the score. I add them as long as they improve my model.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1142955,
      "author_name": "bowaka",
      "author_url": "",
      "post_date": "01/07/2021 17:27:23",
      "content": "<p>Hi ! I will personnaly make a detailed presentation of my current LGB that managed to reach 0.792 LB when the competition is over ! Hope it can help 🙂</p>",
      "votes": null,
      "replies": [
        {
          "id": 1143062,
          "author_name": "ajayhayagreeve",
          "author_url": "",
          "post_date": "01/07/2021 18:32:04",
          "content": "<p>Yeah. Will be very helpful. Thanks<br>\nAlso if you. ensemble with publicly available Transformer model, you will definitely get a boost in score. My 0.005 boost is just  with ensemble of transformer model, publicly available. This one <br>\n<a href=\"https://www.kaggle.com/gilfernandes/riiid-self-attention-transformer\" target=\"_blank\">https://www.kaggle.com/gilfernandes/riiid-self-attention-transformer</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1143020,
      "author_name": "icfstat",
      "author_url": "",
      "post_date": "01/07/2021 18:08:05",
      "content": "<p>I have an lgbm model with 792 in public leaderboard and 797 in local validation. It will be very interesting to read the write-up of lgbm solutions with 0.8+(if that is possible). How far a boosting model can go with this data? What are the\"magic\" features? Any especial treatment in the data?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1143023,
          "author_name": "bowaka",
          "author_url": "",
          "post_date": "01/07/2021 18:09:59",
          "content": "<p>I get same validation and same public ! Curious to see the different write-up too !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1143090,
      "author_name": "yangxiaoshuai",
      "author_url": "",
      "post_date": "01/07/2021 18:44:58",
      "content": "<p>my lgb CV 0.795 LB may be 0.792or0.791, there are about five or four important features I think, and more train data score better.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1143096,
          "author_name": "bowaka",
          "author_url": "",
          "post_date": "01/07/2021 18:51:54",
          "content": "<p>I will be curious to see a model using the top magic features of everybody !</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1143611,
          "author_name": "yangxiaoshuai",
          "author_url": "",
          "post_date": "01/08/2021 01:11:08",
          "content": "<p>I found three magic features at least by top LGBM solutions, all amazing,</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1143129,
      "author_name": "mingpan07",
      "author_url": "",
      "post_date": "01/07/2021 19:05:52",
      "content": "<p>Can confirm a single LGB model trained on 1m rows can reach 0.800 LB. If trained on the full dataset I think it could be competitive for a prize by itself. Unfortunately I didn't have the compute to train on more data in time as I had bugs until this morning and it takes a long time to generate the features. But I'm satisfied with the design even if I couldn't push out the best version in the end. </p>\n<p>Looking forward to reading solutions soon. For me, the \"magic\" was to map the competition data to the open source Ednet data, which has some extra information. It added ~0.01. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1143186,
          "author_name": "rdizzl3",
          "author_url": "",
          "post_date": "01/07/2021 19:32:20",
          "content": "<p>I can confirm this also and for the same reasons + holiday plans I just couldn't find the time to execute my models properly. The engineering alone is a big part of this competition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1143447,
          "author_name": "lrtmonkey",
          "author_url": "",
          "post_date": "01/07/2021 22:54:09",
          "content": "<p>You used the data (Ednet) other than the training data to achieve 0.800?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1143484,
          "author_name": "mingpan07",
          "author_url": "",
          "post_date": "01/07/2021 23:30:18",
          "content": "<p>Yes, the data you can find at <a href=\"https://github.com/riiid/ednet\" target=\"_blank\">https://github.com/riiid/ednet</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1143438,
      "author_name": "beautifulmoment",
      "author_url": "",
      "post_date": "01/07/2021 22:46:56",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5326475%2Fa1a714ace0c8b191c66b3b89b144a3ee%2FFotor_161005881525124.jpg?generation=1610058848758943&amp;alt=media\" alt=\"\"></p>\n<p>Above is my used features, cv about 0.7926.., lb 0.792<br>\nI only use kaggle notebook, using last 29M, sample 15M to train</p>\n<p>upm = user_part_mean<br>\ntb = user_last_samebundle_recency1<br>\ntp = user_last_samepart_recency1<br>\nuccm = user_concent_correctness_mean<br>\ncucm = content_user_correctness_mean<br>\nubcm = user_bundle_correctness_mean<br>\nus = user_correct<br>\nups = user_part_correct</p>\n<p>I am looking forward to see someone's solutions which can achieve 0.795+</p>",
      "votes": null,
      "replies": [
        {
          "id": 1143472,
          "author_name": "fredegrec",
          "author_url": "",
          "post_date": "01/07/2021 23:21:11",
          "content": "<p>I have 0.792 with some another features. Maybe, if add them to your solution it will be 0.795+</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1143944,
      "author_name": "tomooinubushi",
      "author_url": "",
      "post_date": "01/08/2021 06:30:39",
      "content": "<p>From my experience, timestamp_recency_4 and timestamp_recency_5 still have some predictive power and improved the score. I add them as long as they improve my model.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1142907": "Hi, \n   I used the following features for the lgb model, Total (49 features )\n\n**continuous variables :**\n\nuser_correct_mean\n\nquestion_mean\n\ntags_encoded_mean - questions having the same set of tags are grouped together and the correct mean for them. \n\nlecture_tag_seen - number of lectures having tag which is in the question / number of tags in ques.\n\nquestion_tag_seen - sum of number of times explanation seen for the question having tag belonging to the one in the current question / number of tags of the question\n\nuser_question_count\n\nquestion_count\n\ntags_encoded_count\n\n'timestamp_recency_1'\n'timestamp_recency_2'\n'timestamp_recency_3', \n\n'user_prior_question_had_explanation_mean' - mean of number of times explanation seen by user\n\n'user_prior_question_elapsed_time_mean' - mean time the user takes to. answer question\n\nprior_question_elapsed_time\n\nnumber_of_watched_lectures\n\naverage_time_between_interactions - timestamp / number of interactions\n\nlast_correct_recency\n\nlast_incorrect_recency\n\ncontinuous_incorrect\n\ncontinuous_correct\n\nhmean_content_id\n\nhmean_tags_encoded_id\n\nanswered_correctly_rollsum - number of correct_answers for last 10 task container ids\n\nanswered_mean_rollsum  - similar to above but mean\n\nrepeated_content_count \n\nrepeated_content_mean\n\nrepeated_part_count \n\nrepeated_part_mean\n\nbundle_mean\n\nbundle_elapsed_time_mean - avg time taken by all users to answer this bundle\n\nbundle_explanation_mean - avg no of times explanation seen by user for this bundle(btw 0 to 1)\n\nmax_difficulty_crt\n\nmin_difficulty_wrng\n\nuser_correct_std\n\npart_max_difficulty_crt\n\npart_min_difficulty_wrng\n\nnumber_of_tags\n\nuser_tag_mean_accuracy - avg of mean correctness for. each tag in in. the question\n\nuser_tag_max_accuracy\n\nuser_tag_min_accuracy\n\nuser_rating_mean\nuser_rating_std\nquestion_rating_mean\nquestion_rating_std\ndraw_prob - All these belong to true_skill\n\nlecture_recency - time interval between the current question and the last seen lecture (1e12 if not seen)\n\nquestion_tag_lecture_max_recency - time interval between the each tag of the current question and the last seen lecture with the same tag. Then take the min of those (1e12 if not seen)\n\nquestion_tag_lecture_min_recency - similar to above but taking max. \n\n\n**Categorical:**\n\nquestion_part\n\nI updated only after completing a bundle. \n\nI have seen in the discussion people with fewer features than mine have reached 0.79+. But I am getting only. 0.784 with these features though I dont think none of these are noise. \n\nI got validation score of 0.785 with this. I have submitted two submissions prior and I have noted that the lb score is atleast 0.005 more than valid score. (I use different validation set - custom i.e Taking last 5% of data for each users as test and plus 2.5% of the whole users as new)\n\nMy last submission with features upto repeated_content_mean ( top 26 features mentioned above ) scored about 0.781. But I dont know where did I go wrong. Any help would be highly appreciated. \n\nAlthough I dont expect reply before this competition ends, but knowing the reason will help me a lot moving forward. Thanks.",
    "1142955": "Hi ! I will personnaly make a detailed presentation of my current LGB that managed to reach 0.792 LB when the competition is over ! Hope it can help 🙂",
    "1143020": "I have an lgbm model with 792 in public leaderboard and 797 in local validation. It will be very interesting to read the write-up of lgbm solutions with 0.8+(if that is possible). How far a boosting model can go with this data? What are the\"magic\" features? Any especial treatment in the data?",
    "1143023": "I get same validation and same public ! Curious to see the different write-up too !",
    "1143062": "Yeah. Will be very helpful. Thanks\nAlso if you. ensemble with publicly available Transformer model, you will definitely get a boost in score. My 0.005 boost is just  with ensemble of transformer model, publicly available. This one \nhttps://www.kaggle.com/gilfernandes/riiid-self-attention-transformer",
    "1143090": "my lgb CV 0.795 LB may be 0.792or0.791, there are about five or four important features I think, and more train data score better.",
    "1143096": "I will be curious to see a model using the top magic features of everybody !",
    "1143129": "Can confirm a single LGB model trained on 1m rows can reach 0.800 LB. If trained on the full dataset I think it could be competitive for a prize by itself. Unfortunately I didn't have the compute to train on more data in time as I had bugs until this morning and it takes a long time to generate the features. But I'm satisfied with the design even if I couldn't push out the best version in the end. \n\nLooking forward to reading solutions soon. For me, the \"magic\" was to map the competition data to the open source Ednet data, which has some extra information. It added ~0.01.",
    "1143186": "I can confirm this also and for the same reasons + holiday plans I just couldn't find the time to execute my models properly. The engineering alone is a big part of this competition.",
    "1143438": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5326475%2Fa1a714ace0c8b191c66b3b89b144a3ee%2FFotor_161005881525124.jpg?generation=1610058848758943&alt=media)\n\nAbove is my used features, cv about 0.7926.., lb 0.792\nI only use kaggle notebook, using last 29M, sample 15M to train\n\nupm = user_part_mean\ntb = user_last_samebundle_recency1\ntp = user_last_samepart_recency1\nuccm = user_concent_correctness_mean\ncucm = content_user_correctness_mean\nubcm = user_bundle_correctness_mean\nus = user_correct\nups = user_part_correct\n\nI am looking forward to see someone's solutions which can achieve 0.795+",
    "1143447": "You used the data (Ednet) other than the training data to achieve 0.800?",
    "1143472": "I have 0.792 with some another features. Maybe, if add them to your solution it will be 0.795+",
    "1143484": "Yes, the data you can find at https://github.com/riiid/ednet",
    "1143611": "I found three magic features at least by top LGBM solutions, all amazing,",
    "1143944": "From my experience, timestamp_recency_4 and timestamp_recency_5 still have some predictive power and improved the score. I add them as long as they improve my model."
  },
  "source": "meta"
}