{
  "id": 54516,
  "title": "Duplicated rows in train and test",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/54516",
  "author_name": "Meta Learner",
  "post_date": "2018-04-14T07:27:51.133000",
  "votes": 0,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I recently found duplicated rows in train.csv and test.csv files. </p>\n\n<blockquote>\n  <p>select ip, click_time, app, channel, device, os, is_attributed,\n  count(*) as total_count </p>\n  \n  <p>from train</p>\n  \n  <p>group by ip, click_time, app, channel, device, os, is_attributed</p>\n  \n  <p>having total_count  &gt;1</p>\n</blockquote>\n\n<p>![Duplicate rows in train.csv][1]</p>\n\n<p>To prevent over fitting while training these need to be filtered.</p>",
  "messages": [
    {
      "id": 313969,
      "postDate": "2018-04-14T07:27:51.133Z",
      "content": "<p>I recently found duplicated rows in train.csv and test.csv files. </p>\n\n<blockquote>\n  <p>select ip, click_time, app, channel, device, os, is_attributed,\n  count(*) as total_count </p>\n  \n  <p>from train</p>\n  \n  <p>group by ip, click_time, app, channel, device, os, is_attributed</p>\n  \n  <p>having total_count  &gt;1</p>\n</blockquote>\n\n<p>![Duplicate rows in train.csv][1]</p>\n\n<p>To prevent over fitting while training these need to be filtered.</p>",
      "rawMarkdown": "I recently found duplicated rows in train.csv and test.csv files. \n\n\n&gt; select ip, click_time, app, channel, device, os, is_attributed,\n&gt; count(*) as total_count \n&gt; \n&gt; from train\n&gt; \n&gt; group by ip, click_time, app, channel, device, os, is_attributed\n&gt; \n&gt; having total_count  &gt;1\n\n![Duplicate rows in train.csv][1]\n\nTo prevent over fitting while training these need to be filtered.\n\n\n"
    },
    {
      "id": 314023,
      "postDate": "2018-04-14T12:22:34.647Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 314023,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-14T12:22:34.647000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "313969": "I recently found duplicated rows in train.csv and test.csv files. \n\n\n&gt; select ip, click_time, app, channel, device, os, is_attributed,\n&gt; count(*) as total_count \n&gt; \n&gt; from train\n&gt; \n&gt; group by ip, click_time, app, channel, device, os, is_attributed\n&gt; \n&gt; having total_count  &gt;1\n\n![Duplicate rows in train.csv][1]\n\nTo prevent over fitting while training these need to be filtered.\n\n\n",
    "314023": ""
  }
}