{
  "id": 193365,
  "title": "[EDA] How to Handle Outlier Students?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/193365",
  "author_name": "",
  "post_date": "2020-10-26T17:42:57.243991500Z",
  "votes": 37,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I took a look at how average correctness per student evolves over time:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4426199%2F84c4f0a9e6a2811da1a0ba41a849def2%2FScreen%20Shot%202020-10-26%20at%206.24.36%20PM.png?generation=1603733843142895&amp;alt=media\" alt=\"\"></p>\n<p>And I found one particular student that seems to stand out from the crowd:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4426199%2Fb175d9e0d71841364ac8a3f1686c9657%2FScreen%20Shot%202020-10-26%20at%206.25.48%20PM.png?generation=1603733906843311&amp;alt=media\" alt=\"\"></p>\n<p>It seems like the student went through the <em>entire curriculum</em> of learning tasks within 2-3 days after starting:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4426199%2F5ad686614279d0e785ded718c40cea1d%2FScreen%20Shot%202020-10-26%20at%206.26.35%20PM.png?generation=1603733924190052&amp;alt=media\" alt=\"\"></p>\n<p>And the average correctness was around 0.27, indicating that the person kept randomly clicking one of the 4 multiple choice options without even reading the question.</p>\n<p>Would it make sense to remove such students from the dataset before training the model?</p>\n<p>The student's ID is <code>7171715</code> for those interested.</p>",
  "messages": [
    {
      "id": "1061025",
      "postDate": "10/26/2020 17:42:57",
      "content": "<p>I took a look at how average correctness per student evolves over time:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4426199%2F84c4f0a9e6a2811da1a0ba41a849def2%2FScreen%20Shot%202020-10-26%20at%206.24.36%20PM.png?generation=1603733843142895&amp;alt=media\" alt=\"\"></p>\n<p>And I found one particular student that seems to stand out from the crowd:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4426199%2Fb175d9e0d71841364ac8a3f1686c9657%2FScreen%20Shot%202020-10-26%20at%206.25.48%20PM.png?generation=1603733906843311&amp;alt=media\" alt=\"\"></p>\n<p>It seems like the student went through the <em>entire curriculum</em> of learning tasks within 2-3 days after starting:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4426199%2F5ad686614279d0e785ded718c40cea1d%2FScreen%20Shot%202020-10-26%20at%206.26.35%20PM.png?generation=1603733924190052&amp;alt=media\" alt=\"\"></p>\n<p>And the average correctness was around 0.27, indicating that the person kept randomly clicking one of the 4 multiple choice options without even reading the question.</p>\n<p>Would it make sense to remove such students from the dataset before training the model?</p>\n<p>The student's ID is <code>7171715</code> for those interested.</p>",
      "rawMarkdown": "I took a look at how average correctness per student evolves over time:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4426199%2F84c4f0a9e6a2811da1a0ba41a849def2%2FScreen%20Shot%202020-10-26%20at%206.24.36%20PM.png?generation=1603733843142895&alt=media)\n\nAnd I found one particular student that seems to stand out from the crowd:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4426199%2Fb175d9e0d71841364ac8a3f1686c9657%2FScreen%20Shot%202020-10-26%20at%206.25.48%20PM.png?generation=1603733906843311&alt=media)\n\nIt seems like the student went through the _entire curriculum_ of learning tasks within 2-3 days after starting:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4426199%2F5ad686614279d0e785ded718c40cea1d%2FScreen%20Shot%202020-10-26%20at%206.26.35%20PM.png?generation=1603733924190052&alt=media)\n\nAnd the average correctness was around 0.27, indicating that the person kept randomly clicking one of the 4 multiple choice options without even reading the question.\n\nWould it make sense to remove such students from the dataset before training the model?\n\nThe student's ID is `7171715` for those interested.",
      "votes": null
    },
    {
      "id": "1061069",
      "postDate": "10/26/2020 18:06:58",
      "content": "<p>Beautiful plots! Which module is it by the way? Doesn't seem like plotly/bokeh to me. Ty for the info!</p>",
      "rawMarkdown": "Beautiful plots! Which module is it by the way? Doesn't seem like plotly/bokeh to me. Ty for the info!",
      "votes": null
    },
    {
      "id": "1061077",
      "postDate": "10/26/2020 18:12:41",
      "content": "<p>These were generated with Tableau.</p>",
      "rawMarkdown": "These were generated with Tableau.",
      "votes": null
    },
    {
      "id": "1061143",
      "postDate": "10/26/2020 19:12:57",
      "content": "<p>Why not to predict 0.27 all the time for him? 😄</p>",
      "rawMarkdown": "Why not to predict 0.27 all the time for him? 😄",
      "votes": null
    },
    {
      "id": "1064126",
      "postDate": "10/29/2020 18:34:46",
      "content": "<p>Some students just have too much free time 😄</p>",
      "rawMarkdown": "Some students just have too much free time 😄",
      "votes": null
    },
    {
      "id": "1092178",
      "postDate": "11/26/2020 15:26:46",
      "content": "<p>If the outliers are very few in number compared to the total data size, your models shouldn't be affected too much. Do be careful when there are multiple outliers!</p>",
      "rawMarkdown": "If the outliers are very few in number compared to the total data size, your models shouldn't be affected too much. Do be careful when there are multiple outliers!",
      "votes": null
    },
    {
      "id": "1092210",
      "postDate": "11/26/2020 15:57:16",
      "content": "<p>Interesting finding.  What if the user_id 7171715 appears in the test set？ I think it should be remained.</p>",
      "rawMarkdown": "Interesting finding.  What if the user_id 7171715 appears in the test set？ I think it should be remained.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1061069,
      "author_name": "adityaecdrid",
      "author_url": "",
      "post_date": "10/26/2020 18:06:58",
      "content": "<p>Beautiful plots! Which module is it by the way? Doesn't seem like plotly/bokeh to me. Ty for the info!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1061077,
          "author_name": "nyakaggle",
          "author_url": "",
          "post_date": "10/26/2020 18:12:41",
          "content": "<p>These were generated with Tableau.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1061143,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "10/26/2020 19:12:57",
          "content": "<p>Why not to predict 0.27 all the time for him? 😄</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1064126,
      "author_name": "rafiko1",
      "author_url": "",
      "post_date": "10/29/2020 18:34:46",
      "content": "<p>Some students just have too much free time 😄</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1092178,
      "author_name": "namansood",
      "author_url": "",
      "post_date": "11/26/2020 15:26:46",
      "content": "<p>If the outliers are very few in number compared to the total data size, your models shouldn't be affected too much. Do be careful when there are multiple outliers!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1092210,
      "author_name": "kazekage",
      "author_url": "",
      "post_date": "11/26/2020 15:57:16",
      "content": "<p>Interesting finding.  What if the user_id 7171715 appears in the test set？ I think it should be remained.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1061025": "I took a look at how average correctness per student evolves over time:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4426199%2F84c4f0a9e6a2811da1a0ba41a849def2%2FScreen%20Shot%202020-10-26%20at%206.24.36%20PM.png?generation=1603733843142895&alt=media)\n\nAnd I found one particular student that seems to stand out from the crowd:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4426199%2Fb175d9e0d71841364ac8a3f1686c9657%2FScreen%20Shot%202020-10-26%20at%206.25.48%20PM.png?generation=1603733906843311&alt=media)\n\nIt seems like the student went through the _entire curriculum_ of learning tasks within 2-3 days after starting:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4426199%2F5ad686614279d0e785ded718c40cea1d%2FScreen%20Shot%202020-10-26%20at%206.26.35%20PM.png?generation=1603733924190052&alt=media)\n\nAnd the average correctness was around 0.27, indicating that the person kept randomly clicking one of the 4 multiple choice options without even reading the question.\n\nWould it make sense to remove such students from the dataset before training the model?\n\nThe student's ID is `7171715` for those interested.",
    "1061069": "Beautiful plots! Which module is it by the way? Doesn't seem like plotly/bokeh to me. Ty for the info!",
    "1061077": "These were generated with Tableau.",
    "1061143": "Why not to predict 0.27 all the time for him? 😄",
    "1064126": "Some students just have too much free time 😄",
    "1092178": "If the outliers are very few in number compared to the total data size, your models shouldn't be affected too much. Do be careful when there are multiple outliers!",
    "1092210": "Interesting finding.  What if the user_id 7171715 appears in the test set？ I think it should be remained."
  },
  "source": "meta"
}