{
  "id": 205324,
  "title": "How to drop the duplicate rows from the train dataset ?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/205324",
  "author_name": "",
  "post_date": "2020-12-19T14:39:33.840202100Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I was performing basic EDA on the train.cv data and found some duplicate rows <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4678376%2Fd213741ea40ac3447766f85164252ac6%2FScreenshot%20(126).png?generation=1608388622774030&amp;alt=media\" alt=\"\"><br>\nand <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4678376%2F29e1e792995509390cdb1d998016c50b%2FScreenshot%20(125).png?generation=1608388718110637&amp;alt=media\" alt=\"\"><br>\nBut not able to deduplicate these rows<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4678376%2F7ee55292096ac23520476563b9469006%2FScreenshot%20(127).png?generation=1608388766048059&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "1118937",
      "postDate": "12/19/2020 14:39:33",
      "content": "<p>I was performing basic EDA on the train.cv data and found some duplicate rows <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4678376%2Fd213741ea40ac3447766f85164252ac6%2FScreenshot%20(126).png?generation=1608388622774030&amp;alt=media\" alt=\"\"><br>\nand <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4678376%2F29e1e792995509390cdb1d998016c50b%2FScreenshot%20(125).png?generation=1608388718110637&amp;alt=media\" alt=\"\"><br>\nBut not able to deduplicate these rows<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4678376%2F7ee55292096ac23520476563b9469006%2FScreenshot%20(127).png?generation=1608388766048059&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I was performing basic EDA on the train.cv data and found some duplicate rows \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4678376%2Fd213741ea40ac3447766f85164252ac6%2FScreenshot%20(126).png?generation=1608388622774030&alt=media)\nand \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4678376%2F29e1e792995509390cdb1d998016c50b%2FScreenshot%20(125).png?generation=1608388718110637&alt=media)\nBut not able to deduplicate these rows\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4678376%2F7ee55292096ac23520476563b9469006%2FScreenshot%20(127).png?generation=1608388766048059&alt=media)",
      "votes": null
    },
    {
      "id": "1119000",
      "postDate": "12/19/2020 16:10:38",
      "content": "<p>the row_id column is different, you need to drop that column first if you want to use your function.</p>",
      "rawMarkdown": "the row_id column is different, you need to drop that column first if you want to use your function.",
      "votes": null
    },
    {
      "id": "1119062",
      "postDate": "12/19/2020 17:24:32",
      "content": "<p>Yes, I did that also but when I run this code </p>\n<h1>Deduplication of entries</h1>\n<p>final= data.drop_duplicates(keep='first', inplace=False)<br>\nfinal.shape <br>\nDue to the large dataset, kernel stops while running &amp;  restarted again </p>",
      "rawMarkdown": "Yes, I did that also but when I run this code \n#Deduplication of entries\nfinal= data.drop_duplicates(keep='first', inplace=False)\nfinal.shape \nDue to the large dataset, kernel stops while running &  restarted again",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1119000,
      "author_name": "bowaka",
      "author_url": "",
      "post_date": "12/19/2020 16:10:38",
      "content": "<p>the row_id column is different, you need to drop that column first if you want to use your function.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1119062,
      "author_name": "paritoshmahto",
      "author_url": "",
      "post_date": "12/19/2020 17:24:32",
      "content": "<p>Yes, I did that also but when I run this code </p>\n<h1>Deduplication of entries</h1>\n<p>final= data.drop_duplicates(keep='first', inplace=False)<br>\nfinal.shape <br>\nDue to the large dataset, kernel stops while running &amp;  restarted again </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1118937": "I was performing basic EDA on the train.cv data and found some duplicate rows \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4678376%2Fd213741ea40ac3447766f85164252ac6%2FScreenshot%20(126).png?generation=1608388622774030&alt=media)\nand \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4678376%2F29e1e792995509390cdb1d998016c50b%2FScreenshot%20(125).png?generation=1608388718110637&alt=media)\nBut not able to deduplicate these rows\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4678376%2F7ee55292096ac23520476563b9469006%2FScreenshot%20(127).png?generation=1608388766048059&alt=media)",
    "1119000": "the row_id column is different, you need to drop that column first if you want to use your function.",
    "1119062": "Yes, I did that also but when I run this code \n#Deduplication of entries\nfinal= data.drop_duplicates(keep='first', inplace=False)\nfinal.shape \nDue to the large dataset, kernel stops while running &  restarted again"
  },
  "source": "meta"
}