{
  "id": 153171,
  "title": "Sequence data fix",
  "url": "/competitions/iwildcam-2020-fgvc7/discussion/153171",
  "author_name": "",
  "post_date": "2020-05-23T13:09:54.712990300Z",
  "votes": 4,
  "comment_count": 1,
  "views": 0,
  "content": "<p>FYI, the sequence (seq_id) labels provided in both the train and test sets are very noisy. Both sets have sequences with tens of thousands of unrelated images. If you relabel the sequences based on their proximity by date and time taken for each location you'll get much better results, assuming you're using sequence information. I saw a 10% accuracy bump from doing this alone.</p>\n\n<p>Thought I'd share as this is an issue with the data and not really due to better classification</p>",
  "messages": [
    {
      "id": "858412",
      "postDate": "05/23/2020 13:09:54",
      "content": "<p>FYI, the sequence (seq_id) labels provided in both the train and test sets are very noisy. Both sets have sequences with tens of thousands of unrelated images. If you relabel the sequences based on their proximity by date and time taken for each location you'll get much better results, assuming you're using sequence information. I saw a 10% accuracy bump from doing this alone.</p>\n\n<p>Thought I'd share as this is an issue with the data and not really due to better classification</p>",
      "rawMarkdown": "FYI, the sequence (seq_id) labels provided in both the train and test sets are very noisy. Both sets have sequences with tens of thousands of unrelated images. If you relabel the sequences based on their proximity by date and time taken for each location you'll get much better results, assuming you're using sequence information. I saw a 10% accuracy bump from doing this alone.\n\nThought I'd share as this is an issue with the data and not really due to better classification",
      "votes": null
    },
    {
      "id": "858675",
      "postDate": "05/23/2020 17:23:50",
      "content": "<p>Good catch, thanks for flagging this! We think those sequence labels are original (from the organization providing the data) but we'll check into it just in case.</p>",
      "rawMarkdown": "Good catch, thanks for flagging this! We think those sequence labels are original (from the organization providing the data) but we'll check into it just in case.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 858675,
      "author_name": "elijahcole",
      "author_url": "",
      "post_date": "05/23/2020 17:23:50",
      "content": "<p>Good catch, thanks for flagging this! We think those sequence labels are original (from the organization providing the data) but we'll check into it just in case.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "858412": "FYI, the sequence (seq_id) labels provided in both the train and test sets are very noisy. Both sets have sequences with tens of thousands of unrelated images. If you relabel the sequences based on their proximity by date and time taken for each location you'll get much better results, assuming you're using sequence information. I saw a 10% accuracy bump from doing this alone.\n\nThought I'd share as this is an issue with the data and not really due to better classification",
    "858675": "Good catch, thanks for flagging this! We think those sequence labels are original (from the organization providing the data) but we'll check into it just in case."
  },
  "source": "meta"
}