{
  "id": 197276,
  "title": "Modelling of randomness and maximal possible roc score",
  "url": "/competitions/riiid-test-answer-prediction/discussion/197276",
  "author_name": "Eugen Keil",
  "post_date": "2020-11-15T13:29:20.406000",
  "votes": 4,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I was wondering whether other people have included the fact in their models that for each question the probability of guessing the correct answer is at least 25% (33% for questions in part 2).</p>\n<p>I was/am trying to build a model that would predict the correct probability that the student will answer correctly and therefore applied linear transformations to the average number of correct answers for each question and user to account for the guessing aspect. To my big surprise this had almost no effect on my score. After looking up how the roc score is defined it now makes sense (it is invariant to such transformations).</p>\n<p>Could someone with more experience in the area shed light as to why we would use the roc score instead of evaluating how good the predictions are directly by using probability theory?</p>\n<p>Playing around with a simple linear distribution question difficulty it seems like the upper bound for the score will be pretty close to the current leaderboard and SAINT+ benchmarks.</p>\n<pre><code>import random\nfrom sklearn.metrics import roc_auc_score\nN = 100000\npredictions = [(i % 10) / 10 for i in range(N)]\nanswers = [1 if random.random() &lt; (i % 10) / 10 else 0 for i in range(N)]\nroc_auc_score(answers, predictions)\n</code></pre>\n<p>returns a score of about 0.83 and the predictions here are best possible.</p>\n<p>Even with perfect knowledge of user skill and all context information present the score seems to be bounded by 0.9 due to the above mentioned fact of 25% chance of guessing correctly.</p>\n<p>Did anyone try to make a better random model that is closer to the actual distribution to determine what the true theoretical upper bound for the roc score for this problem is?</p>",
  "messages": [
    {
      "id": 1078967,
      "postDate": "2020-11-15T13:29:20.407Z",
      "content": "<p>I was wondering whether other people have included the fact in their models that for each question the probability of guessing the correct answer is at least 25% (33% for questions in part 2).</p>\n<p>I was/am trying to build a model that would predict the correct probability that the student will answer correctly and therefore applied linear transformations to the average number of correct answers for each question and user to account for the guessing aspect. To my big surprise this had almost no effect on my score. After looking up how the roc score is defined it now makes sense (it is invariant to such transformations).</p>\n<p>Could someone with more experience in the area shed light as to why we would use the roc score instead of evaluating how good the predictions are directly by using probability theory?</p>\n<p>Playing around with a simple linear distribution question difficulty it seems like the upper bound for the score will be pretty close to the current leaderboard and SAINT+ benchmarks.</p>\n<pre><code>import random\nfrom sklearn.metrics import roc_auc_score\nN = 100000\npredictions = [(i % 10) / 10 for i in range(N)]\nanswers = [1 if random.random() &lt; (i % 10) / 10 else 0 for i in range(N)]\nroc_auc_score(answers, predictions)\n</code></pre>\n<p>returns a score of about 0.83 and the predictions here are best possible.</p>\n<p>Even with perfect knowledge of user skill and all context information present the score seems to be bounded by 0.9 due to the above mentioned fact of 25% chance of guessing correctly.</p>\n<p>Did anyone try to make a better random model that is closer to the actual distribution to determine what the true theoretical upper bound for the roc score for this problem is?</p>",
      "rawMarkdown": "I was wondering whether other people have included the fact in their models that for each question the probability of guessing the correct answer is at least 25% (33% for questions in part 2).\n\nI was/am trying to build a model that would predict the correct probability that the student will answer correctly and therefore applied linear transformations to the average number of correct answers for each question and user to account for the guessing aspect. To my big surprise this had almost no effect on my score. After looking up how the roc score is defined it now makes sense (it is invariant to such transformations).\n\nCould someone with more experience in the area shed light as to why we would use the roc score instead of evaluating how good the predictions are directly by using probability theory?\n\nPlaying around with a simple linear distribution question difficulty it seems like the upper bound for the score will be pretty close to the current leaderboard and SAINT+ benchmarks.\n```\n\nimport random\nfrom sklearn.metrics import roc_auc_score\nN = 100000\npredictions = [(i % 10) / 10 for i in range(N)]\nanswers = [1 if random.random() < (i % 10) / 10 else 0 for i in range(N)]\nroc_auc_score(answers, predictions)\n```\n\nreturns a score of about 0.83 and the predictions here are best possible.\n\nEven with perfect knowledge of user skill and all context information present the score seems to be bounded by 0.9 due to the above mentioned fact of 25% chance of guessing correctly.\n\nDid anyone try to make a better random model that is closer to the actual distribution to determine what the true theoretical upper bound for the roc score for this problem is?",
      "votes": 4
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1078967": "I was wondering whether other people have included the fact in their models that for each question the probability of guessing the correct answer is at least 25% (33% for questions in part 2).\n\nI was/am trying to build a model that would predict the correct probability that the student will answer correctly and therefore applied linear transformations to the average number of correct answers for each question and user to account for the guessing aspect. To my big surprise this had almost no effect on my score. After looking up how the roc score is defined it now makes sense (it is invariant to such transformations).\n\nCould someone with more experience in the area shed light as to why we would use the roc score instead of evaluating how good the predictions are directly by using probability theory?\n\nPlaying around with a simple linear distribution question difficulty it seems like the upper bound for the score will be pretty close to the current leaderboard and SAINT+ benchmarks.\n```\n\nimport random\nfrom sklearn.metrics import roc_auc_score\nN = 100000\npredictions = [(i % 10) / 10 for i in range(N)]\nanswers = [1 if random.random() < (i % 10) / 10 else 0 for i in range(N)]\nroc_auc_score(answers, predictions)\n```\n\nreturns a score of about 0.83 and the predictions here are best possible.\n\nEven with perfect knowledge of user skill and all context information present the score seems to be bounded by 0.9 due to the above mentioned fact of 25% chance of guessing correctly.\n\nDid anyone try to make a better random model that is closer to the actual distribution to determine what the true theoretical upper bound for the roc score for this problem is?"
  }
}