{
  "id": 199235,
  "title": "[Data bug?] DUPLICATED `content_id` on the same user's task?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/199235",
  "author_name": "",
  "post_date": "2020-11-25T01:01:11.259829100Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I found somewhat strange behavior:</p>\n<pre><code>df = pd.read_csv(\"train.csv\")\ncond1 = df['user_id'] == 1477452754\ncond2 = df['task_container_id'] == 8\ndf[cond1 &amp; cond2]\n&gt;&gt;&gt;\n\n                          row_id  timestamp     user_id       content_id  content_type_id  \\\n69450655  69450655     429864  1477452754         152                0\n69450656  69450656     429864  1477452754         152                0\n69450657  69450657     429864  1477452754         152                0\n69450658  69450658     429864  1477452754         152                0\n\n          task_container_id  user_answer  answered_correctly  \\\n69450655                  8            1                   1\n69450656                  8            1                   1\n69450657                  8            1                   1\n69450658                  8            0                   0\n\n                    prior_question_elapsed_time         prior_question_had_explanation\n69450655                     7000.000                          False\n69450656                     7000.000                          False\n69450657                     7000.000                          False\n69450658                     7000.000                          False\n</code></pre>\n<p>As you can see above, one user solved specific <code>content_id</code> multiple times in a single same task.</p>\n<p>It might be like he or she had failed to pass the question by giving the wrong answer so tried multiple times but data showed that he got the correct answers 3 times.</p>\n<p>I found this kind of data (total 393 rows). You can find them by:</p>\n<pre><code>cond =  train_df[['user_id', 'task_container_id', 'content_type_id', `content_id`]].duplicated(keep=False)\ntrain_df[cond]\n</code></pre>\n<p>Any opinion on this situation?</p>",
  "messages": [
    {
      "id": "1089989",
      "postDate": "11/25/2020 01:01:11",
      "content": "<p>I found somewhat strange behavior:</p>\n<pre><code>df = pd.read_csv(\"train.csv\")\ncond1 = df['user_id'] == 1477452754\ncond2 = df['task_container_id'] == 8\ndf[cond1 &amp; cond2]\n&gt;&gt;&gt;\n\n                          row_id  timestamp     user_id       content_id  content_type_id  \\\n69450655  69450655     429864  1477452754         152                0\n69450656  69450656     429864  1477452754         152                0\n69450657  69450657     429864  1477452754         152                0\n69450658  69450658     429864  1477452754         152                0\n\n          task_container_id  user_answer  answered_correctly  \\\n69450655                  8            1                   1\n69450656                  8            1                   1\n69450657                  8            1                   1\n69450658                  8            0                   0\n\n                    prior_question_elapsed_time         prior_question_had_explanation\n69450655                     7000.000                          False\n69450656                     7000.000                          False\n69450657                     7000.000                          False\n69450658                     7000.000                          False\n</code></pre>\n<p>As you can see above, one user solved specific <code>content_id</code> multiple times in a single same task.</p>\n<p>It might be like he or she had failed to pass the question by giving the wrong answer so tried multiple times but data showed that he got the correct answers 3 times.</p>\n<p>I found this kind of data (total 393 rows). You can find them by:</p>\n<pre><code>cond =  train_df[['user_id', 'task_container_id', 'content_type_id', `content_id`]].duplicated(keep=False)\ntrain_df[cond]\n</code></pre>\n<p>Any opinion on this situation?</p>",
      "rawMarkdown": "I found somewhat strange behavior:\n\n```\ndf = pd.read_csv(\"train.csv\")\ncond1 = df['user_id'] == 1477452754\ncond2 = df['task_container_id'] == 8\ndf[cond1 & cond2]\n>>>\n\n                          row_id  timestamp     user_id       content_id  content_type_id  \\\n69450655  69450655     429864  1477452754         152                0\n69450656  69450656     429864  1477452754         152                0\n69450657  69450657     429864  1477452754         152                0\n69450658  69450658     429864  1477452754         152                0\n\n          task_container_id  user_answer  answered_correctly  \\\n69450655                  8            1                   1\n69450656                  8            1                   1\n69450657                  8            1                   1\n69450658                  8            0                   0\n\n                    prior_question_elapsed_time         prior_question_had_explanation\n69450655                     7000.000                          False\n69450656                     7000.000                          False\n69450657                     7000.000                          False\n69450658                     7000.000                          False\n\n```\n\nAs you can see above, one user solved specific `content_id` multiple times in a single same task.\n\nIt might be like he or she had failed to pass the question by giving the wrong answer so tried multiple times but data showed that he got the correct answers 3 times.\n\nI found this kind of data (total 393 rows). You can find them by:\n\n```\ncond =  train_df[['user_id', 'task_container_id', 'content_type_id', `content_id`]].duplicated(keep=False)\ntrain_df[cond]\n```\n\n\nAny opinion on this situation?",
      "votes": null
    },
    {
      "id": "1090140",
      "postDate": "11/25/2020 05:42:37",
      "content": "<p>Thanks for sharing, It's so strange, and timestamps remain same. </p>",
      "rawMarkdown": "Thanks for sharing, It's so strange, and timestamps remain same.",
      "votes": null
    },
    {
      "id": "1093414",
      "postDate": "11/27/2020 17:37:09",
      "content": "<p>Huge datasets will always include errors.  Data cleaning likely missed the issue you have identified.</p>\n<p>What to do is the time honored issue!!!</p>\n<p>I would vote to delete the 393 rows - but ignore is not a bad second thought.</p>",
      "rawMarkdown": "Huge datasets will always include errors.  Data cleaning likely missed the issue you have identified.\n\nWhat to do is the time honored issue!!!\n\nI would vote to delete the 393 rows - but ignore is not a bad second thought.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1090140,
      "author_name": "kazekage",
      "author_url": "",
      "post_date": "11/25/2020 05:42:37",
      "content": "<p>Thanks for sharing, It's so strange, and timestamps remain same. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1093414,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "11/27/2020 17:37:09",
      "content": "<p>Huge datasets will always include errors.  Data cleaning likely missed the issue you have identified.</p>\n<p>What to do is the time honored issue!!!</p>\n<p>I would vote to delete the 393 rows - but ignore is not a bad second thought.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1089989": "I found somewhat strange behavior:\n\n```\ndf = pd.read_csv(\"train.csv\")\ncond1 = df['user_id'] == 1477452754\ncond2 = df['task_container_id'] == 8\ndf[cond1 & cond2]\n>>>\n\n                          row_id  timestamp     user_id       content_id  content_type_id  \\\n69450655  69450655     429864  1477452754         152                0\n69450656  69450656     429864  1477452754         152                0\n69450657  69450657     429864  1477452754         152                0\n69450658  69450658     429864  1477452754         152                0\n\n          task_container_id  user_answer  answered_correctly  \\\n69450655                  8            1                   1\n69450656                  8            1                   1\n69450657                  8            1                   1\n69450658                  8            0                   0\n\n                    prior_question_elapsed_time         prior_question_had_explanation\n69450655                     7000.000                          False\n69450656                     7000.000                          False\n69450657                     7000.000                          False\n69450658                     7000.000                          False\n\n```\n\nAs you can see above, one user solved specific `content_id` multiple times in a single same task.\n\nIt might be like he or she had failed to pass the question by giving the wrong answer so tried multiple times but data showed that he got the correct answers 3 times.\n\nI found this kind of data (total 393 rows). You can find them by:\n\n```\ncond =  train_df[['user_id', 'task_container_id', 'content_type_id', `content_id`]].duplicated(keep=False)\ntrain_df[cond]\n```\n\n\nAny opinion on this situation?",
    "1090140": "Thanks for sharing, It's so strange, and timestamps remain same.",
    "1093414": "Huge datasets will always include errors.  Data cleaning likely missed the issue you have identified.\n\nWhat to do is the time honored issue!!!\n\nI would vote to delete the 393 rows - but ignore is not a bad second thought."
  },
  "source": "meta"
}