{
  "id": 227680,
  "title": "Duplicates in train.csv file",
  "url": "/competitions/hotel-id-2021-fgvc8/discussion/227680",
  "author_name": "Shanmukh",
  "post_date": "2021-03-21T18:29:28.791000",
  "votes": 8,
  "comment_count": 0,
  "views": 0,
  "content": "<p>There are <strong><code>2</code></strong> duplicates in <strong>train.csv</strong> file. Make sure you remove them before training. You can use the following code snippet.</p>\n<p><code>train_df = train_df.drop_duplicates(subset=['image'])</code></p>\n<p>Number of images in <strong>train folder</strong> = <code>97554</code><br>\nNumber of images in <strong>train.csv</strong> file = <code>97556</code></p>\n<p>Hope this helps.</p>",
  "messages": [
    {
      "id": 1247462,
      "postDate": "2021-03-21T18:29:28.793Z",
      "content": "<p>There are <strong><code>2</code></strong> duplicates in <strong>train.csv</strong> file. Make sure you remove them before training. You can use the following code snippet.</p>\n<p><code>train_df = train_df.drop_duplicates(subset=['image'])</code></p>\n<p>Number of images in <strong>train folder</strong> = <code>97554</code><br>\nNumber of images in <strong>train.csv</strong> file = <code>97556</code></p>\n<p>Hope this helps.</p>",
      "rawMarkdown": "There are **`2`** duplicates in **train.csv** file. Make sure you remove them before training. You can use the following code snippet.\n\n`train_df = train_df.drop_duplicates(subset=['image'])`\n\nNumber of images in **train folder** = `97554`\nNumber of images in **train.csv** file = `97556`\n\nHope this helps.",
      "votes": 8
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1247462": "There are **`2`** duplicates in **train.csv** file. Make sure you remove them before training. You can use the following code snippet.\n\n`train_df = train_df.drop_duplicates(subset=['image'])`\n\nNumber of images in **train folder** = `97554`\nNumber of images in **train.csv** file = `97556`\n\nHope this helps."
  }
}