{
  "id": 208244,
  "title": "Users guessing the answer correct/wrong based on 'correct_answer' trend",
  "url": "/competitions/riiid-test-answer-prediction/discussion/208244",
  "author_name": "",
  "post_date": "2021-01-02T13:58:20.580071400Z",
  "votes": 3,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hey there all~ </p>\n<p>I've only just started working on my EDA (way too late actually XD) and just when I was thinking about things I could do  I remembered the times I took Multiple Choice Question exams. I remembered that students get a bit uneasy when they choose the same LETTER answer (in our case [ 3, 2, 1, 0 ] numbers) multiple times in a row. And because of that we tend to guess the answer based on that trend  when we are not sure or mark the incorrect the answer because it feels like a trap set up by the exam makers. So, I tested if my above statement was correct by following steps:</p>\n<ul>\n<li>I only worked on 10M rows of the training data</li>\n<li>Found consecutive questions (or content_id, question_id) on each users (user_id)</li>\n<li>And from that checked if correct_answer had at least 3 same values in a row<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5845653%2F3d0b08620b0f8a9f3d8329cfb7d553cb%2FScreen%20Shot%202021-01-02%20at%2021.23.42.png?generation=1609595335852893&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5845653%2Ffb5ff5a592c5d1fd89bc279e388cf59f%2FScreen%20Shot%202021-01-02%20at%2021.22.02.png?generation=1609595402043305&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5845653%2F454ac35106fc09b72a4cefa956dff2f8%2FScreen%20Shot%202021-01-02%20at%2021.17.23.png?generation=1609595420452670&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5845653%2F59d4e988997b7255e0b766368cd48475%2FScreen%20Shot%202021-01-02%20at%2021.16.29.png?generation=1609595439205662&amp;alt=media\" alt=\"\"></li>\n</ul>\n<p>As you can see above some users had excellent streak of correct answers while some had failed in between. In this 10M rows of data about 20 users had consecutive questions so I think in the full training set there will be a lot more. And as there will not be new question in the testing set I think there will also be consecutive questions with same number correct answer. As I'm beginner I'm not sure how to proceed with this finding. Perhaps this finding is not useful at all and may only add more noise to our prediction. What do you think?</p>\n<p>P.S As this is my first discussion I must have worded what I wanted to inform slightly off. <br>\nCheers…</p>",
  "messages": [
    {
      "id": "1135788",
      "postDate": "01/02/2021 13:58:20",
      "content": "<p>Hey there all~ </p>\n<p>I've only just started working on my EDA (way too late actually XD) and just when I was thinking about things I could do  I remembered the times I took Multiple Choice Question exams. I remembered that students get a bit uneasy when they choose the same LETTER answer (in our case [ 3, 2, 1, 0 ] numbers) multiple times in a row. And because of that we tend to guess the answer based on that trend  when we are not sure or mark the incorrect the answer because it feels like a trap set up by the exam makers. So, I tested if my above statement was correct by following steps:</p>\n<ul>\n<li>I only worked on 10M rows of the training data</li>\n<li>Found consecutive questions (or content_id, question_id) on each users (user_id)</li>\n<li>And from that checked if correct_answer had at least 3 same values in a row<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5845653%2F3d0b08620b0f8a9f3d8329cfb7d553cb%2FScreen%20Shot%202021-01-02%20at%2021.23.42.png?generation=1609595335852893&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5845653%2Ffb5ff5a592c5d1fd89bc279e388cf59f%2FScreen%20Shot%202021-01-02%20at%2021.22.02.png?generation=1609595402043305&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5845653%2F454ac35106fc09b72a4cefa956dff2f8%2FScreen%20Shot%202021-01-02%20at%2021.17.23.png?generation=1609595420452670&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5845653%2F59d4e988997b7255e0b766368cd48475%2FScreen%20Shot%202021-01-02%20at%2021.16.29.png?generation=1609595439205662&amp;alt=media\" alt=\"\"></li>\n</ul>\n<p>As you can see above some users had excellent streak of correct answers while some had failed in between. In this 10M rows of data about 20 users had consecutive questions so I think in the full training set there will be a lot more. And as there will not be new question in the testing set I think there will also be consecutive questions with same number correct answer. As I'm beginner I'm not sure how to proceed with this finding. Perhaps this finding is not useful at all and may only add more noise to our prediction. What do you think?</p>\n<p>P.S As this is my first discussion I must have worded what I wanted to inform slightly off. <br>\nCheers…</p>",
      "rawMarkdown": "Hey there all~ \n\nI've only just started working on my EDA (way too late actually XD) and just when I was thinking about things I could do  I remembered the times I took Multiple Choice Question exams. I remembered that students get a bit uneasy when they choose the same LETTER answer (in our case [ 3, 2, 1, 0 ] numbers) multiple times in a row. And because of that we tend to guess the answer based on that trend  when we are not sure or mark the incorrect the answer because it feels like a trap set up by the exam makers. So, I tested if my above statement was correct by following steps:\n- I only worked on 10M rows of the training data\n- Found consecutive questions (or content_id, question_id) on each users (user_id)\n- And from that checked if correct_answer had at least 3 same values in a row\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5845653%2F3d0b08620b0f8a9f3d8329cfb7d553cb%2FScreen%20Shot%202021-01-02%20at%2021.23.42.png?generation=1609595335852893&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5845653%2Ffb5ff5a592c5d1fd89bc279e388cf59f%2FScreen%20Shot%202021-01-02%20at%2021.22.02.png?generation=1609595402043305&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5845653%2F454ac35106fc09b72a4cefa956dff2f8%2FScreen%20Shot%202021-01-02%20at%2021.17.23.png?generation=1609595420452670&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5845653%2F59d4e988997b7255e0b766368cd48475%2FScreen%20Shot%202021-01-02%20at%2021.16.29.png?generation=1609595439205662&alt=media)\n\nAs you can see above some users had excellent streak of correct answers while some had failed in between. In this 10M rows of data about 20 users had consecutive questions so I think in the full training set there will be a lot more. And as there will not be new question in the testing set I think there will also be consecutive questions with same number correct answer. As I'm beginner I'm not sure how to proceed with this finding. Perhaps this finding is not useful at all and may only add more noise to our prediction. What do you think?\n\nP.S As this is my first discussion I must have worded what I wanted to inform slightly off. \nCheers...",
      "votes": null
    },
    {
      "id": "1135878",
      "postDate": "01/02/2021 15:13:26",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/temuujinerdene\" target=\"_blank\">@temuujinerdene</a>,</p>\n<p>Welcome to the competition and perhaps more importantly, welcome to the participating side of the discussions :-).</p>\n<p>The points you bring up are very valid. However, although it hasn't been demonstratively proven and the host hasn't released any information confirming this, I think it's safe to say that the order the questions were presented to the users (A-B-C-D) vs the order we have in the dataset are not one and the same.</p>\n<p>Each question in the dataset we're given has only a single correct_response for all users. Whereas the ITS (Intelligent Training System) app this data was collected within is called <a href=\"https://www.riiid.co/en/product\" target=\"_blank\">SANTA</a>, and it'd be absolutely <strong>wild</strong> to assume they didn't do either a random shuffling of the response placement, or even intelligently permuting them to address the very concern you raise here:</p>\n<blockquote>\n  <p>students get a bit uneasy when they choose the same LETTER answer (in our case [ 3, 2, 1, 0 ] numbers) multiple times in a row</p>\n</blockquote>\n<p>If we were given truly raw dhat, this sort of EDA would surely provide some signal. But in our case I personally do not believe there is anything here.</p>",
      "rawMarkdown": "Hi @temuujinerdene,\n\nWelcome to the competition and perhaps more importantly, welcome to the participating side of the discussions :-).\n\nThe points you bring up are very valid. However, although it hasn't been demonstratively proven and the host hasn't released any information confirming this, I think it's safe to say that the order the questions were presented to the users (A-B-C-D) vs the order we have in the dataset are not one and the same.\n\nEach question in the dataset we're given has only a single correct_response for all users. Whereas the ITS (Intelligent Training System) app this data was collected within is called [SANTA](https://www.riiid.co/en/product), and it'd be absolutely **wild** to assume they didn't do either a random shuffling of the response placement, or even intelligently permuting them to address the very concern you raise here:\n\n> students get a bit uneasy when they choose the same LETTER answer (in our case [ 3, 2, 1, 0 ] numbers) multiple times in a row\n\nIf we were given truly raw dhat, this sort of EDA would surely provide some signal. But in our case I personally do not believe there is anything here.",
      "votes": null
    },
    {
      "id": "1135892",
      "postDate": "01/02/2021 15:22:20",
      "content": "<p>Your idea is nice and i did check on this around 3 weeks back, A plot below for the user_id \"1660941992\", the person as correctness streak as high as 35-40.</p>\n<blockquote>\n  <p>welcome to the participating side of the discussions :-).</p>\n</blockquote>\n<p>Happy to see this line <a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a>! </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2F6c5d7ca0c33eda48dac9b8374a8dbf8f%2FScreenshot%202021-01-02%20at%208.51.03%20PM.png?generation=1609600885181634&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Your idea is nice and i did check on this around 3 weeks back, A plot below for the user_id \"1660941992\", the person as correctness streak as high as 35-40.\n\n> welcome to the participating side of the discussions :-).\n\nHappy to see this line @authman! \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2F6c5d7ca0c33eda48dac9b8374a8dbf8f%2FScreenshot%202021-01-02%20at%208.51.03%20PM.png?generation=1609600885181634&alt=media)",
      "votes": null
    },
    {
      "id": "1136164",
      "postDate": "01/02/2021 19:13:22",
      "content": "<p>I also checked this and confirmed that it can increase auc when working with the training set. But I am struggling to understand how to implement this with the actual test set…</p>",
      "rawMarkdown": "I also checked this and confirmed that it can increase auc when working with the training set. But I am struggling to understand how to implement this with the actual test set...",
      "votes": null
    },
    {
      "id": "1136210",
      "postDate": "01/02/2021 20:38:20",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/michaelj8\" target=\"_blank\">@michaelj8</a></p>\n<blockquote>\n  <p>confirmed that it can increase auc when working with the training set</p>\n</blockquote>\n<p>How did you confirm it increased auc in training set without that transferring into test set? By that I mean, the only way I am aware of to confirm would be to split train set and test it. But if such a test were possible, the same test should be replicateable against our submission set.</p>",
      "rawMarkdown": "Hi @michaelj8\n\n> confirmed that it can increase auc when working with the training set\n\nHow did you confirm it increased auc in training set without that transferring into test set? By that I mean, the only way I am aware of to confirm would be to split train set and test it. But if such a test were possible, the same test should be replicateable against our submission set.",
      "votes": null
    },
    {
      "id": "1136330",
      "postDate": "01/03/2021 00:53:12",
      "content": "<p>I see, thank you very much for your response.</p>",
      "rawMarkdown": "I see, thank you very much for your response.",
      "votes": null
    },
    {
      "id": "1140183",
      "postDate": "01/05/2021 20:50:04",
      "content": "<p>In my desperation for FE, I've revisited every single idea I've had along with everything mentioned by others—even things my gut disagreed with. In looking at this feature particularly, it does indeed seem as though there is some signal here. I don't know yet if its generalizeable or not. Completed some initial EDA on it and now attempting to featurize + test it.</p>",
      "rawMarkdown": "In my desperation for FE, I've revisited every single idea I've had along with everything mentioned by others—even things my gut disagreed with. In looking at this feature particularly, it does indeed seem as though there is some signal here. I don't know yet if its generalizeable or not. Completed some initial EDA on it and now attempting to featurize + test it.",
      "votes": null
    },
    {
      "id": "1145200",
      "postDate": "01/09/2021 01:13:56",
      "content": "<p>I used something similar but in two different ways</p>\n<ol>\n<li>I called a variable user answer bias (% user answered any particular number), and I did see some potential pattern during EDA in the those answered questions per user vs the actual proportion of answers asked in the exam-some users would have much more 2 answers than was actually correct in the exam. When I was a student, if i didn't know the answer, I would always select D… for no particular reason, I was just 'bias' to that answer. </li>\n</ol>\n<p>In some of my models it help (as determined by feature importance in a lightgbm model which is what I was building everything off of). But as I started to add different variables this \"user answer bias\" became much less important and I decided to leave it out as it wasn't adding AUC improvement to my model. </p>\n<ol>\n<li>I also did use most recently answered in a row correctly for my model- if your hot your hot ..right?. The interesting thing is I was getting really high AUC (92%) on my training set and decent on MY test set 83-84% (which on the competition test set might have been 79%)… but I didn't have the time to work on getting this to work in the competition test set… I was really thrown off by how different the competition test set was from the training set and didn't leave enough time to code for this correctly. Also, the decision tree looked strange to me, all leaf GINI's were .61 to .68   so I didn't really investigate more because I would feel uneasy if my leaf's had such close GINI's.. I really don't know what that means.</li>\n</ol>",
      "rawMarkdown": "I used something similar but in two different ways\n\n1. I called a variable user answer bias (% user answered any particular number), and I did see some potential pattern during EDA in the those answered questions per user vs the actual proportion of answers asked in the exam-some users would have much more 2 answers than was actually correct in the exam. When I was a student, if i didn't know the answer, I would always select D... for no particular reason, I was just 'bias' to that answer. \n\nIn some of my models it help (as determined by feature importance in a lightgbm model which is what I was building everything off of). But as I started to add different variables this \"user answer bias\" became much less important and I decided to leave it out as it wasn't adding AUC improvement to my model. \n\n2. I also did use most recently answered in a row correctly for my model- if your hot your hot ..right?. The interesting thing is I was getting really high AUC (92%) on my training set and decent on MY test set 83-84% (which on the competition test set might have been 79%)... but I didn't have the time to work on getting this to work in the competition test set... I was really thrown off by how different the competition test set was from the training set and didn't leave enough time to code for this correctly. Also, the decision tree looked strange to me, all leaf GINI's were .61 to .68   so I didn't really investigate more because I would feel uneasy if my leaf's had such close GINI's.. I really don't know what that means.",
      "votes": null
    },
    {
      "id": "1145204",
      "postDate": "01/09/2021 01:17:42",
      "content": "<p>I used MY test set from the training set if that makes sense. I have another reply above, but basically MY test set and the competition test set are different. I didn't really leave enough time to code for the formatting difference in the competition test set… there were some nuisances and I just didn't leave the time to figure it out.. (two kids, full time job… basically I wasn't worried about winning this thing[nor do I think I have the skills yet to do so)</p>",
      "rawMarkdown": "I used MY test set from the training set if that makes sense. I have another reply above, but basically MY test set and the competition test set are different. I didn't really leave enough time to code for the formatting difference in the competition test set... there were some nuisances and I just didn't leave the time to figure it out.. (two kids, full time job... basically I wasn't worried about winning this thing[nor do I think I have the skills yet to do so)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1135878,
      "author_name": "authman",
      "author_url": "",
      "post_date": "01/02/2021 15:13:26",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/temuujinerdene\" target=\"_blank\">@temuujinerdene</a>,</p>\n<p>Welcome to the competition and perhaps more importantly, welcome to the participating side of the discussions :-).</p>\n<p>The points you bring up are very valid. However, although it hasn't been demonstratively proven and the host hasn't released any information confirming this, I think it's safe to say that the order the questions were presented to the users (A-B-C-D) vs the order we have in the dataset are not one and the same.</p>\n<p>Each question in the dataset we're given has only a single correct_response for all users. Whereas the ITS (Intelligent Training System) app this data was collected within is called <a href=\"https://www.riiid.co/en/product\" target=\"_blank\">SANTA</a>, and it'd be absolutely <strong>wild</strong> to assume they didn't do either a random shuffling of the response placement, or even intelligently permuting them to address the very concern you raise here:</p>\n<blockquote>\n  <p>students get a bit uneasy when they choose the same LETTER answer (in our case [ 3, 2, 1, 0 ] numbers) multiple times in a row</p>\n</blockquote>\n<p>If we were given truly raw dhat, this sort of EDA would surely provide some signal. But in our case I personally do not believe there is anything here.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1136330,
          "author_name": "temuujinerdene",
          "author_url": "",
          "post_date": "01/03/2021 00:53:12",
          "content": "<p>I see, thank you very much for your response.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1140183,
          "author_name": "authman",
          "author_url": "",
          "post_date": "01/05/2021 20:50:04",
          "content": "<p>In my desperation for FE, I've revisited every single idea I've had along with everything mentioned by others—even things my gut disagreed with. In looking at this feature particularly, it does indeed seem as though there is some signal here. I don't know yet if its generalizeable or not. Completed some initial EDA on it and now attempting to featurize + test it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1135892,
      "author_name": "adityaecdrid",
      "author_url": "",
      "post_date": "01/02/2021 15:22:20",
      "content": "<p>Your idea is nice and i did check on this around 3 weeks back, A plot below for the user_id \"1660941992\", the person as correctness streak as high as 35-40.</p>\n<blockquote>\n  <p>welcome to the participating side of the discussions :-).</p>\n</blockquote>\n<p>Happy to see this line <a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a>! </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2F6c5d7ca0c33eda48dac9b8374a8dbf8f%2FScreenshot%202021-01-02%20at%208.51.03%20PM.png?generation=1609600885181634&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1136164,
      "author_name": "michaelj8",
      "author_url": "",
      "post_date": "01/02/2021 19:13:22",
      "content": "<p>I also checked this and confirmed that it can increase auc when working with the training set. But I am struggling to understand how to implement this with the actual test set…</p>",
      "votes": null,
      "replies": [
        {
          "id": 1136210,
          "author_name": "authman",
          "author_url": "",
          "post_date": "01/02/2021 20:38:20",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/michaelj8\" target=\"_blank\">@michaelj8</a></p>\n<blockquote>\n  <p>confirmed that it can increase auc when working with the training set</p>\n</blockquote>\n<p>How did you confirm it increased auc in training set without that transferring into test set? By that I mean, the only way I am aware of to confirm would be to split train set and test it. But if such a test were possible, the same test should be replicateable against our submission set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1145204,
          "author_name": "michaelj8",
          "author_url": "",
          "post_date": "01/09/2021 01:17:42",
          "content": "<p>I used MY test set from the training set if that makes sense. I have another reply above, but basically MY test set and the competition test set are different. I didn't really leave enough time to code for the formatting difference in the competition test set… there were some nuisances and I just didn't leave the time to figure it out.. (two kids, full time job… basically I wasn't worried about winning this thing[nor do I think I have the skills yet to do so)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1145200,
      "author_name": "michaelj8",
      "author_url": "",
      "post_date": "01/09/2021 01:13:56",
      "content": "<p>I used something similar but in two different ways</p>\n<ol>\n<li>I called a variable user answer bias (% user answered any particular number), and I did see some potential pattern during EDA in the those answered questions per user vs the actual proportion of answers asked in the exam-some users would have much more 2 answers than was actually correct in the exam. When I was a student, if i didn't know the answer, I would always select D… for no particular reason, I was just 'bias' to that answer. </li>\n</ol>\n<p>In some of my models it help (as determined by feature importance in a lightgbm model which is what I was building everything off of). But as I started to add different variables this \"user answer bias\" became much less important and I decided to leave it out as it wasn't adding AUC improvement to my model. </p>\n<ol>\n<li>I also did use most recently answered in a row correctly for my model- if your hot your hot ..right?. The interesting thing is I was getting really high AUC (92%) on my training set and decent on MY test set 83-84% (which on the competition test set might have been 79%)… but I didn't have the time to work on getting this to work in the competition test set… I was really thrown off by how different the competition test set was from the training set and didn't leave enough time to code for this correctly. Also, the decision tree looked strange to me, all leaf GINI's were .61 to .68   so I didn't really investigate more because I would feel uneasy if my leaf's had such close GINI's.. I really don't know what that means.</li>\n</ol>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1135788": "Hey there all~ \n\nI've only just started working on my EDA (way too late actually XD) and just when I was thinking about things I could do  I remembered the times I took Multiple Choice Question exams. I remembered that students get a bit uneasy when they choose the same LETTER answer (in our case [ 3, 2, 1, 0 ] numbers) multiple times in a row. And because of that we tend to guess the answer based on that trend  when we are not sure or mark the incorrect the answer because it feels like a trap set up by the exam makers. So, I tested if my above statement was correct by following steps:\n- I only worked on 10M rows of the training data\n- Found consecutive questions (or content_id, question_id) on each users (user_id)\n- And from that checked if correct_answer had at least 3 same values in a row\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5845653%2F3d0b08620b0f8a9f3d8329cfb7d553cb%2FScreen%20Shot%202021-01-02%20at%2021.23.42.png?generation=1609595335852893&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5845653%2Ffb5ff5a592c5d1fd89bc279e388cf59f%2FScreen%20Shot%202021-01-02%20at%2021.22.02.png?generation=1609595402043305&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5845653%2F454ac35106fc09b72a4cefa956dff2f8%2FScreen%20Shot%202021-01-02%20at%2021.17.23.png?generation=1609595420452670&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5845653%2F59d4e988997b7255e0b766368cd48475%2FScreen%20Shot%202021-01-02%20at%2021.16.29.png?generation=1609595439205662&alt=media)\n\nAs you can see above some users had excellent streak of correct answers while some had failed in between. In this 10M rows of data about 20 users had consecutive questions so I think in the full training set there will be a lot more. And as there will not be new question in the testing set I think there will also be consecutive questions with same number correct answer. As I'm beginner I'm not sure how to proceed with this finding. Perhaps this finding is not useful at all and may only add more noise to our prediction. What do you think?\n\nP.S As this is my first discussion I must have worded what I wanted to inform slightly off. \nCheers...",
    "1135878": "Hi @temuujinerdene,\n\nWelcome to the competition and perhaps more importantly, welcome to the participating side of the discussions :-).\n\nThe points you bring up are very valid. However, although it hasn't been demonstratively proven and the host hasn't released any information confirming this, I think it's safe to say that the order the questions were presented to the users (A-B-C-D) vs the order we have in the dataset are not one and the same.\n\nEach question in the dataset we're given has only a single correct_response for all users. Whereas the ITS (Intelligent Training System) app this data was collected within is called [SANTA](https://www.riiid.co/en/product), and it'd be absolutely **wild** to assume they didn't do either a random shuffling of the response placement, or even intelligently permuting them to address the very concern you raise here:\n\n> students get a bit uneasy when they choose the same LETTER answer (in our case [ 3, 2, 1, 0 ] numbers) multiple times in a row\n\nIf we were given truly raw dhat, this sort of EDA would surely provide some signal. But in our case I personally do not believe there is anything here.",
    "1135892": "Your idea is nice and i did check on this around 3 weeks back, A plot below for the user_id \"1660941992\", the person as correctness streak as high as 35-40.\n\n> welcome to the participating side of the discussions :-).\n\nHappy to see this line @authman! \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2F6c5d7ca0c33eda48dac9b8374a8dbf8f%2FScreenshot%202021-01-02%20at%208.51.03%20PM.png?generation=1609600885181634&alt=media)",
    "1136164": "I also checked this and confirmed that it can increase auc when working with the training set. But I am struggling to understand how to implement this with the actual test set...",
    "1136210": "Hi @michaelj8\n\n> confirmed that it can increase auc when working with the training set\n\nHow did you confirm it increased auc in training set without that transferring into test set? By that I mean, the only way I am aware of to confirm would be to split train set and test it. But if such a test were possible, the same test should be replicateable against our submission set.",
    "1136330": "I see, thank you very much for your response.",
    "1140183": "In my desperation for FE, I've revisited every single idea I've had along with everything mentioned by others—even things my gut disagreed with. In looking at this feature particularly, it does indeed seem as though there is some signal here. I don't know yet if its generalizeable or not. Completed some initial EDA on it and now attempting to featurize + test it.",
    "1145200": "I used something similar but in two different ways\n\n1. I called a variable user answer bias (% user answered any particular number), and I did see some potential pattern during EDA in the those answered questions per user vs the actual proportion of answers asked in the exam-some users would have much more 2 answers than was actually correct in the exam. When I was a student, if i didn't know the answer, I would always select D... for no particular reason, I was just 'bias' to that answer. \n\nIn some of my models it help (as determined by feature importance in a lightgbm model which is what I was building everything off of). But as I started to add different variables this \"user answer bias\" became much less important and I decided to leave it out as it wasn't adding AUC improvement to my model. \n\n2. I also did use most recently answered in a row correctly for my model- if your hot your hot ..right?. The interesting thing is I was getting really high AUC (92%) on my training set and decent on MY test set 83-84% (which on the competition test set might have been 79%)... but I didn't have the time to work on getting this to work in the competition test set... I was really thrown off by how different the competition test set was from the training set and didn't leave enough time to code for this correctly. Also, the decision tree looked strange to me, all leaf GINI's were .61 to .68   so I didn't really investigate more because I would feel uneasy if my leaf's had such close GINI's.. I really don't know what that means.",
    "1145204": "I used MY test set from the training set if that makes sense. I have another reply above, but basically MY test set and the competition test set are different. I didn't really leave enough time to code for the formatting difference in the competition test set... there were some nuisances and I just didn't leave the time to figure it out.. (two kids, full time job... basically I wasn't worried about winning this thing[nor do I think I have the skills yet to do so)"
  },
  "source": "meta"
}