{
  "id": 74165,
  "title": "ERROR: The value '158928' in the key column 'object_id' has already been defined (Line 15140, Column 316)",
  "url": "/competitions/PLAsTiCC-2018/discussion/74165",
  "author_name": "",
  "post_date": "2018-12-09T15:12:26.729639300Z",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Can someone please explain this error and how can I remove it?</p>",
  "messages": [
    {
      "id": "436096",
      "postDate": "12/09/2018 15:12:26",
      "content": "<p>Can someone please explain this error and how can I remove it?</p>",
      "rawMarkdown": "Can someone please explain this error and how can I remove it?",
      "votes": null
    },
    {
      "id": "436102",
      "postDate": "12/09/2018 15:40:24",
      "content": "<p>158928 object_id, is appearing twice in your submission list, you need to have only one, and find out which other object_id is missing.</p>",
      "rawMarkdown": "158928 object_id, is appearing twice in your submission list, you need to have only one, and find out which other object_id is missing.",
      "votes": null
    },
    {
      "id": "436173",
      "postDate": "12/09/2018 19:31:50",
      "content": "<p>If you are using the test set chunking code that is in many of the public kernels such as this one <a href=\"https://www.kaggle.com/cttsai/smote-on-lgbm-w-ideas-from-kernels-and-discussion/code\">https://www.kaggle.com/cttsai/smote-on-lgbm-w-ideas-from-kernels-and-discussion/code</a> There will be duplicate object_id in the first submission file that is created. In the above kernel the duplicates are removed by grouping by object_id and taking the mean.</p>\n\n<pre><code>z = pd.read_csv(filename)\nprint(\"Shape BEFORE grouping: {}\".format(z.shape))\nz = z.groupby('object_id').mean()\nprint(\"Shape AFTER grouping: {}\".format(z.shape))\nz.to_csv('single_{}'.format(filename), index=True)\n</code></pre>\n\n<p>The file with prefix \"single_\" is the one you want to submit.</p>",
      "rawMarkdown": "If you are using the test set chunking code that is in many of the public kernels such as this one https://www.kaggle.com/cttsai/smote-on-lgbm-w-ideas-from-kernels-and-discussion/code There will be duplicate object_id in the first submission file that is created. In the above kernel the duplicates are removed by grouping by object_id and taking the mean.\n\n    z = pd.read_csv(filename)\n    print(\"Shape BEFORE grouping: {}\".format(z.shape))\n    z = z.groupby('object_id').mean()\n    print(\"Shape AFTER grouping: {}\".format(z.shape))\n    z.to_csv('single_{}'.format(filename), index=True)\n\nThe file with prefix \"single_\" is the one you want to submit.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 436102,
      "author_name": "chinta",
      "author_url": "",
      "post_date": "12/09/2018 15:40:24",
      "content": "<p>158928 object_id, is appearing twice in your submission list, you need to have only one, and find out which other object_id is missing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 436173,
      "author_name": "jackvial",
      "author_url": "",
      "post_date": "12/09/2018 19:31:50",
      "content": "<p>If you are using the test set chunking code that is in many of the public kernels such as this one <a href=\"https://www.kaggle.com/cttsai/smote-on-lgbm-w-ideas-from-kernels-and-discussion/code\">https://www.kaggle.com/cttsai/smote-on-lgbm-w-ideas-from-kernels-and-discussion/code</a> There will be duplicate object_id in the first submission file that is created. In the above kernel the duplicates are removed by grouping by object_id and taking the mean.</p>\n\n<pre><code>z = pd.read_csv(filename)\nprint(\"Shape BEFORE grouping: {}\".format(z.shape))\nz = z.groupby('object_id').mean()\nprint(\"Shape AFTER grouping: {}\".format(z.shape))\nz.to_csv('single_{}'.format(filename), index=True)\n</code></pre>\n\n<p>The file with prefix \"single_\" is the one you want to submit.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "436096": "Can someone please explain this error and how can I remove it?",
    "436102": "158928 object_id, is appearing twice in your submission list, you need to have only one, and find out which other object_id is missing.",
    "436173": "If you are using the test set chunking code that is in many of the public kernels such as this one https://www.kaggle.com/cttsai/smote-on-lgbm-w-ideas-from-kernels-and-discussion/code There will be duplicate object_id in the first submission file that is created. In the above kernel the duplicates are removed by grouping by object_id and taking the mean.\n\n    z = pd.read_csv(filename)\n    print(\"Shape BEFORE grouping: {}\".format(z.shape))\n    z = z.groupby('object_id').mean()\n    print(\"Shape AFTER grouping: {}\".format(z.shape))\n    z.to_csv('single_{}'.format(filename), index=True)\n\nThe file with prefix \"single_\" is the one you want to submit."
  },
  "source": "meta"
}