{
  "id": 209589,
  "title": "An interesting feature.",
  "url": "/competitions/riiid-test-answer-prediction/discussion/209589",
  "author_name": "Shihao Shao",
  "post_date": "2021-01-08T00:38:47.895000",
  "votes": 27,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Although my rank is not so high but just at 155th place, I want to share a powerful feature with you, as I did not see others mention it. It is the mean of mean correctness of content that certain user has done before, which judges the difficulty of tasks that users have done. If an user's correctness is very high, but, with the value mentioned above is also high, it is less reasoned that this user has good skill.</p>\n<p>This feature gives me a rise of about 0.002~0.003.</p>\n<p>If you have some other tricky features, please let me know :D</p>",
  "messages": [
    {
      "id": 1143577,
      "postDate": "2021-01-08T00:38:47.897Z",
      "content": "<p>Although my rank is not so high but just at 155th place, I want to share a powerful feature with you, as I did not see others mention it. It is the mean of mean correctness of content that certain user has done before, which judges the difficulty of tasks that users have done. If an user's correctness is very high, but, with the value mentioned above is also high, it is less reasoned that this user has good skill.</p>\n<p>This feature gives me a rise of about 0.002~0.003.</p>\n<p>If you have some other tricky features, please let me know :D</p>",
      "rawMarkdown": "Although my rank is not so high but just at 155th place, I want to share a powerful feature with you, as I did not see others mention it. It is the mean of mean correctness of content that certain user has done before, which judges the difficulty of tasks that users have done. If an user's correctness is very high, but, with the value mentioned above is also high, it is less reasoned that this user has good skill.\n \nThis feature gives me a rise of about 0.002~0.003.\n\nIf you have some other tricky features, please let me know :D",
      "votes": 27
    },
    {
      "id": 1143646,
      "postDate": "2021-01-08T01:35:22.497Z",
      "content": "<p>I used these features as well and some variants and it was a 0.003 - 0.004 boost on my CV. There was another one that I found that was the cumulative sum of answered correctly and prior question had explanation divided by the cumulative sum of questions answered correctly (shifted by 1 to avoid leakage). This gave me a 0.002 boost on my CV as well</p>",
      "rawMarkdown": "I used these features as well and some variants and it was a 0.003 - 0.004 boost on my CV. There was another one that I found that was the cumulative sum of answered correctly and prior question had explanation divided by the cumulative sum of questions answered correctly (shifted by 1 to avoid leakage). This gave me a 0.002 boost on my CV as well",
      "votes": 3
    },
    {
      "id": 1143854,
      "postDate": "2021-01-08T05:18:05.630Z",
      "content": "<p>Awesome insight!</p>",
      "rawMarkdown": "Awesome insight!",
      "votes": 1
    },
    {
      "id": 1143624,
      "postDate": "2021-01-08T01:18:16.337Z",
      "content": "<p>Yes, this competition is all about magic/tricky features. The most powerful features in my code is lag_time(user`s current timestamp - last timestamp) and user_content_attempted(similar as attemp_no in many public kernels, but only record whether user attempted this question before), both 2 features gave me almost 0.01 boost, which is huge for this comptition. Thanks for sharing and happy kaggling.</p>",
      "rawMarkdown": "Yes, this competition is all about magic/tricky features. The most powerful features in my code is lag_time(user`s current timestamp - last timestamp) and user_content_attempted(similar as attemp_no in many public kernels, but only record whether user attempted this question before), both 2 features gave me almost 0.01 boost, which is huge for this comptition. Thanks for sharing and happy kaggling.",
      "votes": 1,
      "replies": [
        {
          "id": 1143638,
          "postDate": "2021-01-08T01:26:11.043Z",
          "content": "<p>Yes, lag_time is really important also in my code. But logically why it is so powerful?🤔 </p>",
          "rawMarkdown": "Yes, lag_time is really important also in my code. But logically why it is so powerful?🤔 "
        },
        {
          "id": 1144052,
          "postDate": "2021-01-08T07:54:55.830Z",
          "content": "<p>In most cases, if someone can solve a problem with little time, then he's quite good at these questions. In other words, there will be high probability he/she answers correctly next time. Of course, it is better to combine this feature with the recent accuracy of answering questions(Someone just guesses the same result). It is better to do some data analysis.  </p>",
          "rawMarkdown": "In most cases, if someone can solve a problem with little time, then he's quite good at these questions. In other words, there will be high probability he/she answers correctly next time. Of course, it is better to combine this feature with the recent accuracy of answering questions(Someone just guesses the same result). It is better to do some data analysis.  ",
          "votes": 1
        },
        {
          "id": 1144056,
          "postDate": "2021-01-08T07:58:11.567Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1144542,
          "postDate": "2021-01-08T14:20:57.750Z",
          "content": "<p>Yes, i got this. <a href=\"https://www.kaggle.com/longyin2\" target=\"_blank\">@longyin2</a> In that case maybe we should drop those extremely high values, as it might be the time between user logout today and login the next day but not the time solving the task. Is that right?</p>",
          "rawMarkdown": "Yes, i got this. @longyin2 In that case maybe we should drop those extremely high values, as it might be the time between user logout today and login the next day but not the time solving the task. Is that right?"
        }
      ]
    },
    {
      "id": 1144791,
      "postDate": "2021-01-08T17:10:31.457Z",
      "content": "<p>I personally got a nice boost from Bayesian Probabilities for each questions, and the clustering that comes with it. I didn't quantified the gain, but I would say at least 0.003-0.004 if I check my progression in the competition over time</p>",
      "rawMarkdown": "I personally got a nice boost from Bayesian Probabilities for each questions, and the clustering that comes with it. I didn't quantified the gain, but I would say at least 0.003-0.004 if I check my progression in the competition over time",
      "votes": 2
    },
    {
      "id": 1144049,
      "postDate": "2021-01-08T07:49:58.977Z",
      "content": "<p>牛逼，我有一个特征是最近5次比赛的方差和难题（90%+）的成功率， 结合在一起提升了0.003左右。 还有一个是记录u_content 平均值，也能提升0.001. 可惜我以为八点只是提交截止，最后版本未能上传。</p>",
      "rawMarkdown": "牛逼，我有一个特征是最近5次比赛的方差和难题（90%+）的成功率， 结合在一起提升了0.003左右。 还有一个是记录u_content 平均值，也能提升0.001. 可惜我以为八点只是提交截止，最后版本未能上传。",
      "votes": 2,
      "replies": [
        {
          "id": 1144558,
          "postDate": "2021-01-08T14:30:00.037Z",
          "content": "<p>哈哈哈哈，我后面几天GCP的内存实在吃不消，什么新模型都跑不动，啥都干不了，从90名一路掉到155</p>",
          "rawMarkdown": "哈哈哈哈，我后面几天GCP的内存实在吃不消，什么新模型都跑不动，啥都干不了，从90名一路掉到155"
        },
        {
          "id": 1144773,
          "postDate": "2021-01-08T17:01:33.650Z",
          "content": "<p>之后的比赛我也会考虑租用服务器来跑，自己电脑太不够用了哈哈哈。大佬已经很厉害了，未来比赛还要多向你们学习。</p>",
          "rawMarkdown": "之后的比赛我也会考虑租用服务器来跑，自己电脑太不够用了哈哈哈。大佬已经很厉害了，未来比赛还要多向你们学习。"
        }
      ]
    },
    {
      "id": 1143744,
      "postDate": "2021-01-08T03:28:57.173Z",
      "content": "<p>Hi, Thank you for the opportunity to announce.<br>\nI used these features that probably not too many person use. I focused on \"user_answer\" and found some features from it.<br>\n  user_answer_variance  :  Variance of 4 choices per contents.<br>\n  user_answer_var_rate  :  Average the variance per users when the user chooses the correct answer.<br>\n  user_incorrect_answer_variance  :  Variance of 3 choices(incorrect) per contents.<br>\n  user_incorrect_answer_var_rate  :   Average the variance per users when the user chooses the incorrect answer.<br>\nThese features gave me about 0.004 boost.</p>",
      "rawMarkdown": "Hi, Thank you for the opportunity to announce.\nI used these features that probably not too many person use. I focused on \"user_answer\" and found some features from it.\n  user_answer_variance  :  Variance of 4 choices per contents.\n  user_answer_var_rate  :  Average the variance per users when the user chooses the correct answer.\n  user_incorrect_answer_variance  :  Variance of 3 choices(incorrect) per contents.\n  user_incorrect_answer_var_rate  :   Average the variance per users when the user chooses the incorrect answer.\nThese features gave me about 0.004 boost.",
      "votes": 2
    },
    {
      "id": 1143617,
      "postDate": "2021-01-08T01:14:10.750Z",
      "content": "<p>I also use this feature and mean content correctness for part.<br>\nSecond feature also give me improvement by 0.002.<br>\nAnother interesting feature if containers in correct order, give about +0.001</p>",
      "rawMarkdown": "I also use this feature and mean content correctness for part.\nSecond feature also give me improvement by 0.002.\nAnother interesting feature if containers in correct order, give about +0.001",
      "votes": 2,
      "replies": [
        {
          "id": 1143635,
          "postDate": "2021-01-08T01:24:16.183Z",
          "content": "<p>Ohhhhh Why didn't I apply it to 'part'!😂 Nice job bro!</p>",
          "rawMarkdown": "Ohhhhh Why didn't I apply it to 'part'!😂 Nice job bro!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1985692,
      "postDate": "2022-10-13T13:46:51.533Z",
      "content": "<p>wow…thanks again</p>",
      "rawMarkdown": "wow...thanks again"
    }
  ],
  "comments": [
    {
      "id": 1143646,
      "author_name": "RDizzl3",
      "author_url": "",
      "post_date": "2021-01-08T01:35:22.497000",
      "content": "<p>I used these features as well and some variants and it was a 0.003 - 0.004 boost on my CV. There was another one that I found that was the cumulative sum of answered correctly and prior question had explanation divided by the cumulative sum of questions answered correctly (shifted by 1 to avoid leakage). This gave me a 0.002 boost on my CV as well</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1143854,
      "author_name": "HinePo",
      "author_url": "",
      "post_date": "2021-01-08T05:18:05.630000",
      "content": "<p>Awesome insight!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1143624,
      "author_name": "Hao",
      "author_url": "",
      "post_date": "2021-01-08T01:18:16.337000",
      "content": "<p>Yes, this competition is all about magic/tricky features. The most powerful features in my code is lag_time(user`s current timestamp - last timestamp) and user_content_attempted(similar as attemp_no in many public kernels, but only record whether user attempted this question before), both 2 features gave me almost 0.01 boost, which is huge for this comptition. Thanks for sharing and happy kaggling.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1143638,
          "author_name": "Shihao Shao",
          "author_url": "",
          "post_date": "2021-01-08T01:26:11.043000",
          "content": "<p>Yes, lag_time is really important also in my code. But logically why it is so powerful?🤔 </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1144052,
          "author_name": "LongYin/杰少",
          "author_url": "",
          "post_date": "2021-01-08T07:54:55.830000",
          "content": "<p>In most cases, if someone can solve a problem with little time, then he's quite good at these questions. In other words, there will be high probability he/she answers correctly next time. Of course, it is better to combine this feature with the recent accuracy of answering questions(Someone just guesses the same result). It is better to do some data analysis.  </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1144056,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-01-08T07:58:11.567000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1144542,
          "author_name": "Shihao Shao",
          "author_url": "",
          "post_date": "2021-01-08T14:20:57.750000",
          "content": "<p>Yes, i got this. <a href=\"https://www.kaggle.com/longyin2\" target=\"_blank\">@longyin2</a> In that case maybe we should drop those extremely high values, as it might be the time between user logout today and login the next day but not the time solving the task. Is that right?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1144791,
      "author_name": "Jacky",
      "author_url": "",
      "post_date": "2021-01-08T17:10:31.457000",
      "content": "<p>I personally got a nice boost from Bayesian Probabilities for each questions, and the clustering that comes with it. I didn't quantified the gain, but I would say at least 0.003-0.004 if I check my progression in the competition over time</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1144049,
      "author_name": "朴大福",
      "author_url": "",
      "post_date": "2021-01-08T07:49:58.977000",
      "content": "<p>牛逼，我有一个特征是最近5次比赛的方差和难题（90%+）的成功率， 结合在一起提升了0.003左右。 还有一个是记录u_content 平均值，也能提升0.001. 可惜我以为八点只是提交截止，最后版本未能上传。</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1144558,
          "author_name": "Shihao Shao",
          "author_url": "",
          "post_date": "2021-01-08T14:30:00.037000",
          "content": "<p>哈哈哈哈，我后面几天GCP的内存实在吃不消，什么新模型都跑不动，啥都干不了，从90名一路掉到155</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1144773,
          "author_name": "朴大福",
          "author_url": "",
          "post_date": "2021-01-08T17:01:33.650000",
          "content": "<p>之后的比赛我也会考虑租用服务器来跑，自己电脑太不够用了哈哈哈。大佬已经很厉害了，未来比赛还要多向你们学习。</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1143744,
      "author_name": "OYM8012",
      "author_url": "",
      "post_date": "2021-01-08T03:28:57.173000",
      "content": "<p>Hi, Thank you for the opportunity to announce.<br>\nI used these features that probably not too many person use. I focused on \"user_answer\" and found some features from it.<br>\n  user_answer_variance  :  Variance of 4 choices per contents.<br>\n  user_answer_var_rate  :  Average the variance per users when the user chooses the correct answer.<br>\n  user_incorrect_answer_variance  :  Variance of 3 choices(incorrect) per contents.<br>\n  user_incorrect_answer_var_rate  :   Average the variance per users when the user chooses the incorrect answer.<br>\nThese features gave me about 0.004 boost.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1143617,
      "author_name": "Alyona Pasevieva",
      "author_url": "",
      "post_date": "2021-01-08T01:14:10.750000",
      "content": "<p>I also use this feature and mean content correctness for part.<br>\nSecond feature also give me improvement by 0.002.<br>\nAnother interesting feature if containers in correct order, give about +0.001</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1143635,
          "author_name": "Shihao Shao",
          "author_url": "",
          "post_date": "2021-01-08T01:24:16.183000",
          "content": "<p>Ohhhhh Why didn't I apply it to 'part'!😂 Nice job bro!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1985692,
      "author_name": "Isaac Duboc",
      "author_url": "",
      "post_date": "2022-10-13T13:46:51.533000",
      "content": "<p>wow…thanks again</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1143577": "Although my rank is not so high but just at 155th place, I want to share a powerful feature with you, as I did not see others mention it. It is the mean of mean correctness of content that certain user has done before, which judges the difficulty of tasks that users have done. If an user's correctness is very high, but, with the value mentioned above is also high, it is less reasoned that this user has good skill.\n \nThis feature gives me a rise of about 0.002~0.003.\n\nIf you have some other tricky features, please let me know :D",
    "1143646": "I used these features as well and some variants and it was a 0.003 - 0.004 boost on my CV. There was another one that I found that was the cumulative sum of answered correctly and prior question had explanation divided by the cumulative sum of questions answered correctly (shifted by 1 to avoid leakage). This gave me a 0.002 boost on my CV as well",
    "1143854": "Awesome insight!",
    "1143624": "Yes, this competition is all about magic/tricky features. The most powerful features in my code is lag_time(user`s current timestamp - last timestamp) and user_content_attempted(similar as attemp_no in many public kernels, but only record whether user attempted this question before), both 2 features gave me almost 0.01 boost, which is huge for this comptition. Thanks for sharing and happy kaggling.",
    "1144791": "I personally got a nice boost from Bayesian Probabilities for each questions, and the clustering that comes with it. I didn't quantified the gain, but I would say at least 0.003-0.004 if I check my progression in the competition over time",
    "1144049": "牛逼，我有一个特征是最近5次比赛的方差和难题（90%+）的成功率， 结合在一起提升了0.003左右。 还有一个是记录u_content 平均值，也能提升0.001. 可惜我以为八点只是提交截止，最后版本未能上传。",
    "1143744": "Hi, Thank you for the opportunity to announce.\nI used these features that probably not too many person use. I focused on \"user_answer\" and found some features from it.\n  user_answer_variance  :  Variance of 4 choices per contents.\n  user_answer_var_rate  :  Average the variance per users when the user chooses the correct answer.\n  user_incorrect_answer_variance  :  Variance of 3 choices(incorrect) per contents.\n  user_incorrect_answer_var_rate  :   Average the variance per users when the user chooses the incorrect answer.\nThese features gave me about 0.004 boost.",
    "1143617": "I also use this feature and mean content correctness for part.\nSecond feature also give me improvement by 0.002.\nAnother interesting feature if containers in correct order, give about +0.001",
    "1985692": "wow...thanks again"
  }
}