{
  "id": 371916,
  "title": "Duplicates in train dataset",
  "url": "/competitions/otto-recommender-system/discussion/371916",
  "author_name": "",
  "post_date": "2022-12-13T07:08:59.199063100Z",
  "votes": 9,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi everyone, <br>\nI just found that train dataset contains duplicated events (or rows). I have checked multiple version of train dataset (csv, parquet). I don't see anyone mentions on discussions that or drop duplicates in public notebooks. Are the duplicates negligible or everyone just doesn't know about them? 👀<br>\nLength after <code>df.drop_duplicates()</code> is <strong>216,384,937</strong> (was 216,716,096)</p>",
  "messages": [
    {
      "id": "2063662",
      "postDate": "12/13/2022 07:08:59",
      "content": "<p>Hi everyone, <br>\nI just found that train dataset contains duplicated events (or rows). I have checked multiple version of train dataset (csv, parquet). I don't see anyone mentions on discussions that or drop duplicates in public notebooks. Are the duplicates negligible or everyone just doesn't know about them? 👀<br>\nLength after <code>df.drop_duplicates()</code> is <strong>216,384,937</strong> (was 216,716,096)</p>",
      "rawMarkdown": "Hi everyone, \nI just found that train dataset contains duplicated events (or rows). I have checked multiple version of train dataset (csv, parquet). I don't see anyone mentions on discussions that or drop duplicates in public notebooks. Are the duplicates negligible or everyone just doesn't know about them? 👀\nLength after `df.drop_duplicates()` is **216,384,937** (was 216,716,096)",
      "votes": null
    },
    {
      "id": "2077967",
      "postDate": "12/28/2022 02:00:04",
      "content": "<p>I tried drop_duplicates() too, but the len is still 216,716,096?</p>",
      "rawMarkdown": "I tried drop_duplicates() too, but the len is still 216,716,096?",
      "votes": null
    },
    {
      "id": "2078673",
      "postDate": "12/28/2022 14:19:15",
      "content": "<p>Hi, which data did you use?</p>",
      "rawMarkdown": "Hi, which data did you use?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2077967,
      "author_name": "juweichan",
      "author_url": "",
      "post_date": "12/28/2022 02:00:04",
      "content": "<p>I tried drop_duplicates() too, but the len is still 216,716,096?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2078673,
          "author_name": "quangphm",
          "author_url": "",
          "post_date": "12/28/2022 14:19:15",
          "content": "<p>Hi, which data did you use?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2063662": "Hi everyone, \nI just found that train dataset contains duplicated events (or rows). I have checked multiple version of train dataset (csv, parquet). I don't see anyone mentions on discussions that or drop duplicates in public notebooks. Are the duplicates negligible or everyone just doesn't know about them? 👀\nLength after `df.drop_duplicates()` is **216,384,937** (was 216,716,096)",
    "2077967": "I tried drop_duplicates() too, but the len is still 216,716,096?",
    "2078673": "Hi, which data did you use?"
  },
  "source": "meta"
}