{
  "id": 52752,
  "title": "Duplicate rows in train, test & test-supplement datasets - how come?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/52752",
  "author_name": "",
  "post_date": "2018-03-22T19:36:33.533148900Z",
  "votes": 7,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Train distinct count (based on \"ip\",\"app\",\"is_attributed\", \"device\",\"os\",\"channel\",\"attributed_time\" ,\"click_time\") is 181,284,636, the raw <code>wc -l train.csv</code> count is 184,903,891 - so that's nearly 3.6 million exact duplicates or about 2% of the rows in train.csv</p>\n\n<p>I've noticed similar proportion (~1.7%) of duplicate rows in test &amp; test-supplement datasets (of course, you would have to exclude click-id to find the duplicates).\nTest: 18,482,750 (non-dups) vs 18,790,469\nTest-supplement: 56,536,838 (non-dups) vs 57,537,505.</p>\n\n<p>Is this an issue with the resolution of the timestamps provided or a miss in the data preparation stage?</p>",
  "messages": [
    {
      "id": "301462",
      "postDate": "03/22/2018 19:36:33",
      "content": "<p>Train distinct count (based on \"ip\",\"app\",\"is_attributed\", \"device\",\"os\",\"channel\",\"attributed_time\" ,\"click_time\") is 181,284,636, the raw <code>wc -l train.csv</code> count is 184,903,891 - so that's nearly 3.6 million exact duplicates or about 2% of the rows in train.csv</p>\n\n<p>I've noticed similar proportion (~1.7%) of duplicate rows in test &amp; test-supplement datasets (of course, you would have to exclude click-id to find the duplicates).\nTest: 18,482,750 (non-dups) vs 18,790,469\nTest-supplement: 56,536,838 (non-dups) vs 57,537,505.</p>\n\n<p>Is this an issue with the resolution of the timestamps provided or a miss in the data preparation stage?</p>",
      "rawMarkdown": "Train distinct count (based on \"ip\",\"app\",\"is_attributed\", \"device\",\"os\",\"channel\",\"attributed_time\" ,\"click_time\") is 181,284,636, the raw `wc -l train.csv` count is 184,903,891 - so that's nearly 3.6 million exact duplicates or about 2% of the rows in train.csv\n\nI've noticed similar proportion (~1.7%) of duplicate rows in test &amp; test-supplement datasets (of course, you would have to exclude click-id to find the duplicates).\nTest: 18,482,750 (non-dups) vs 18,790,469\nTest-supplement: 56,536,838 (non-dups) vs 57,537,505.\n\nIs this an issue with the resolution of the timestamps provided or a miss in the data preparation stage?",
      "votes": null
    },
    {
      "id": "301746",
      "postDate": "03/23/2018 07:27:16",
      "content": "<p>I hate cleaning. Can you break that down into how many had <code>is_attributed=0</code> vs 1? We can probably ignore the<code>is_attributed=0</code> duplicates. (Presumably <code>is_attributed</code> did not exhibit both 0 and 1 for the exact same variable combination?)</p>\n\n<p>If that is correct data, then that sounds like crappy malware producing multiple clicks within a second.</p>\n\n<p>Anyway can we get clarification on this?</p>",
      "rawMarkdown": "I hate cleaning. Can you break that down into how many had `is_attributed=0` vs 1? We can probably ignore the`is_attributed=0` duplicates. (Presumably `is_attributed` did not exhibit both 0 and 1 for the exact same variable combination?)\n\nIf that is correct data, then that sounds like crappy malware producing multiple clicks within a second.\n\nAnyway can we get clarification on this?",
      "votes": null
    },
    {
      "id": "302123",
      "postDate": "03/23/2018 17:56:36",
      "content": "<p>This could contribute to count features working so well, if these are bot-generated. All of them capture these duplicates.</p>\n\n<p>On the other hand, it could just be a logging issue as well, depending on how they aggregate their logs.</p>",
      "rawMarkdown": "This could contribute to count features working so well, if these are bot-generated. All of them capture these duplicates.\n\nOn the other hand, it could just be a logging issue as well, depending on how they aggregate their logs.",
      "votes": null
    },
    {
      "id": "302183",
      "postDate": "03/23/2018 19:17:21",
      "content": "<p>If duplicates are present in both train and test, then I do not think we should remove duplicates from the train set but try including them in validation set.</p>",
      "rawMarkdown": "If duplicates are present in both train and test, then I do not think we should remove duplicates from the train set but try including them in validation set.",
      "votes": null
    },
    {
      "id": "302435",
      "postDate": "03/24/2018 03:38:22",
      "content": "<p>Maybe the timestamps we are getting are rounded to the nearest seconds, while original bot frequency was at a nanosecond level?  How else to explain same ip/device/etc clicking at the same time?</p>",
      "rawMarkdown": "Maybe the timestamps we are getting are rounded to the nearest seconds, while original bot frequency was at a nanosecond level?  How else to explain same ip/device/etc clicking at the same time?",
      "votes": null
    },
    {
      "id": "302482",
      "postDate": "03/24/2018 06:13:19",
      "content": "<p>Yes, that's the one way to think about it.</p>",
      "rawMarkdown": "Yes, that's the one way to think about it.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 301746,
      "author_name": "smcinerney",
      "author_url": "",
      "post_date": "03/23/2018 07:27:16",
      "content": "<p>I hate cleaning. Can you break that down into how many had <code>is_attributed=0</code> vs 1? We can probably ignore the<code>is_attributed=0</code> duplicates. (Presumably <code>is_attributed</code> did not exhibit both 0 and 1 for the exact same variable combination?)</p>\n\n<p>If that is correct data, then that sounds like crappy malware producing multiple clicks within a second.</p>\n\n<p>Anyway can we get clarification on this?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 302123,
      "author_name": "anttip",
      "author_url": "",
      "post_date": "03/23/2018 17:56:36",
      "content": "<p>This could contribute to count features working so well, if these are bot-generated. All of them capture these duplicates.</p>\n\n<p>On the other hand, it could just be a logging issue as well, depending on how they aggregate their logs.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 302183,
      "author_name": "sohaibomar",
      "author_url": "",
      "post_date": "03/23/2018 19:17:21",
      "content": "<p>If duplicates are present in both train and test, then I do not think we should remove duplicates from the train set but try including them in validation set.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 302435,
      "author_name": "yuliagm",
      "author_url": "",
      "post_date": "03/24/2018 03:38:22",
      "content": "<p>Maybe the timestamps we are getting are rounded to the nearest seconds, while original bot frequency was at a nanosecond level?  How else to explain same ip/device/etc clicking at the same time?</p>",
      "votes": null,
      "replies": [
        {
          "id": 302482,
          "author_name": "sohaibomar",
          "author_url": "",
          "post_date": "03/24/2018 06:13:19",
          "content": "<p>Yes, that's the one way to think about it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "301462": "Train distinct count (based on \"ip\",\"app\",\"is_attributed\", \"device\",\"os\",\"channel\",\"attributed_time\" ,\"click_time\") is 181,284,636, the raw `wc -l train.csv` count is 184,903,891 - so that's nearly 3.6 million exact duplicates or about 2% of the rows in train.csv\n\nI've noticed similar proportion (~1.7%) of duplicate rows in test &amp; test-supplement datasets (of course, you would have to exclude click-id to find the duplicates).\nTest: 18,482,750 (non-dups) vs 18,790,469\nTest-supplement: 56,536,838 (non-dups) vs 57,537,505.\n\nIs this an issue with the resolution of the timestamps provided or a miss in the data preparation stage?",
    "301746": "I hate cleaning. Can you break that down into how many had `is_attributed=0` vs 1? We can probably ignore the`is_attributed=0` duplicates. (Presumably `is_attributed` did not exhibit both 0 and 1 for the exact same variable combination?)\n\nIf that is correct data, then that sounds like crappy malware producing multiple clicks within a second.\n\nAnyway can we get clarification on this?",
    "302123": "This could contribute to count features working so well, if these are bot-generated. All of them capture these duplicates.\n\nOn the other hand, it could just be a logging issue as well, depending on how they aggregate their logs.",
    "302183": "If duplicates are present in both train and test, then I do not think we should remove duplicates from the train set but try including them in validation set.",
    "302435": "Maybe the timestamps we are getting are rounded to the nearest seconds, while original bot frequency was at a nanosecond level?  How else to explain same ip/device/etc clicking at the same time?",
    "302482": "Yes, that's the one way to think about it."
  },
  "source": "meta"
}