{
  "id": 33032,
  "title": "Trouble With Training Set Extraction",
  "url": "/competitions/noaa-fisheries-steller-sea-lion-population-count/discussion/33032",
  "author_name": "",
  "post_date": "2017-05-15T19:52:22.054899Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hello,</p>\n\n<p>I am having some trouble extracting the training data from the large downloadable file. I am interested in the \"TrainDotted\" and \"Train\" folders. </p>\n\n<p>When I inspect the images contained within i get: \nError interpreting JPEG image file (Not a JPEG file: starts with 0xda 0x5a)</p>\n\n<p>Here is what I have tried extracting with:</p>\n\n<p>Linux - ArchLinux - 7zip \nLinux - ArchLinux - Engrampa Archive Manager\nWindows 7 - Winrar (Also on a separate computer)</p>\n\n<p>In Linux I do get a successful extraction, but then the images are corrupted in some way as described further above. Changing the filename to \".png\" does not fix this.</p>\n\n<p>In Windows the extraction simply fails, giving a \"CRC failed in -imagepath-. The file is corrupt\" error message.</p>\n\n<p>I haven't found any similar comments on here, which i find strange since it seems to be a computer independent issue. Help would be appreciated, need me some dataz.</p>",
  "messages": [
    {
      "id": "182777",
      "postDate": "05/15/2017 19:52:22",
      "content": "<p>Hello,</p>\n\n<p>I am having some trouble extracting the training data from the large downloadable file. I am interested in the \"TrainDotted\" and \"Train\" folders. </p>\n\n<p>When I inspect the images contained within i get: \nError interpreting JPEG image file (Not a JPEG file: starts with 0xda 0x5a)</p>\n\n<p>Here is what I have tried extracting with:</p>\n\n<p>Linux - ArchLinux - 7zip \nLinux - ArchLinux - Engrampa Archive Manager\nWindows 7 - Winrar (Also on a separate computer)</p>\n\n<p>In Linux I do get a successful extraction, but then the images are corrupted in some way as described further above. Changing the filename to \".png\" does not fix this.</p>\n\n<p>In Windows the extraction simply fails, giving a \"CRC failed in -imagepath-. The file is corrupt\" error message.</p>\n\n<p>I haven't found any similar comments on here, which i find strange since it seems to be a computer independent issue. Help would be appreciated, need me some dataz.</p>",
      "rawMarkdown": "Hello,\n\nI am having some trouble extracting the training data from the large downloadable file. I am interested in the \"TrainDotted\" and \"Train\" folders. \n\nWhen I inspect the images contained within i get: \nError interpreting JPEG image file (Not a JPEG file: starts with 0xda 0x5a)\n\nHere is what I have tried extracting with:\n\nLinux - ArchLinux - 7zip \nLinux - ArchLinux - Engrampa Archive Manager\nWindows 7 - Winrar (Also on a separate computer)\n\nIn Linux I do get a successful extraction, but then the images are corrupted in some way as described further above. Changing the filename to \".png\" does not fix this.\n\nIn Windows the extraction simply fails, giving a \"CRC failed in -imagepath-. The file is corrupt\" error message.\n\nI haven't found any similar comments on here, which i find strange since it seems to be a computer independent issue. Help would be appreciated, need me some dataz.",
      "votes": null
    },
    {
      "id": "183525",
      "postDate": "05/18/2017 13:19:16",
      "content": "<p>It sounds like your downloaded archive is corrupt, did you download directly? It might have failed early. </p>\n\n<p>Try downloading via the torrent, each chuck is validated as it's downloaded.</p>",
      "rawMarkdown": "It sounds like your downloaded archive is corrupt, did you download directly? It might have failed early. \n\nTry downloading via the torrent, each chuck is validated as it's downloaded.",
      "votes": null
    },
    {
      "id": "183537",
      "postDate": "05/18/2017 13:54:45",
      "content": "<p>Is there a link to the torrent? I may be missing something. And yes, I only tried the direct download.</p>",
      "rawMarkdown": "Is there a link to the torrent? I may be missing something. And yes, I only tried the direct download.",
      "votes": null
    },
    {
      "id": "183543",
      "postDate": "05/18/2017 14:13:35",
      "content": "<p>Yeah it's the 3rd file down on the data page. Download the .torrent file then open it in something like <a href=\"http://www.utorrent.com/\">uTorrent</a> to download the data.</p>\n\n<p>I can't see any checksum values on the download page, but the MD5 hash of my (working) copy is: 09630241b306c6fb2093133bcb1319a8</p>",
      "rawMarkdown": "Yeah it's the 3rd file down on the data page. Download the .torrent file then open it in something like [uTorrent][1] to download the data.\n\nI can't see any checksum values on the download page, but the MD5 hash of my (working) copy is: 09630241b306c6fb2093133bcb1319a8\n\n\n  [1]: http://www.utorrent.com/",
      "votes": null
    },
    {
      "id": "183545",
      "postDate": "05/18/2017 14:15:52",
      "content": "<p>Wow, I had totally ignored that. Thanks for pointing it out. I will go ahead and try the torrent and see what happens.</p>",
      "rawMarkdown": "Wow, I had totally ignored that. Thanks for pointing it out. I will go ahead and try the torrent and see what happens.",
      "votes": null
    },
    {
      "id": "183933",
      "postDate": "05/19/2017 19:53:07",
      "content": "<p>It still seems to be happening with the torrent on my Linux system. Any thoughts?</p>",
      "rawMarkdown": "It still seems to be happening with the torrent on my Linux system. Any thoughts?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 183525,
      "author_name": "garethjns",
      "author_url": "",
      "post_date": "05/18/2017 13:19:16",
      "content": "<p>It sounds like your downloaded archive is corrupt, did you download directly? It might have failed early. </p>\n\n<p>Try downloading via the torrent, each chuck is validated as it's downloaded.</p>",
      "votes": null,
      "replies": [
        {
          "id": 183537,
          "author_name": "saiguysci",
          "author_url": "",
          "post_date": "05/18/2017 13:54:45",
          "content": "<p>Is there a link to the torrent? I may be missing something. And yes, I only tried the direct download.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 183543,
          "author_name": "garethjns",
          "author_url": "",
          "post_date": "05/18/2017 14:13:35",
          "content": "<p>Yeah it's the 3rd file down on the data page. Download the .torrent file then open it in something like <a href=\"http://www.utorrent.com/\">uTorrent</a> to download the data.</p>\n\n<p>I can't see any checksum values on the download page, but the MD5 hash of my (working) copy is: 09630241b306c6fb2093133bcb1319a8</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 183545,
          "author_name": "saiguysci",
          "author_url": "",
          "post_date": "05/18/2017 14:15:52",
          "content": "<p>Wow, I had totally ignored that. Thanks for pointing it out. I will go ahead and try the torrent and see what happens.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 183933,
          "author_name": "saiguysci",
          "author_url": "",
          "post_date": "05/19/2017 19:53:07",
          "content": "<p>It still seems to be happening with the torrent on my Linux system. Any thoughts?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "182777": "Hello,\n\nI am having some trouble extracting the training data from the large downloadable file. I am interested in the \"TrainDotted\" and \"Train\" folders. \n\nWhen I inspect the images contained within i get: \nError interpreting JPEG image file (Not a JPEG file: starts with 0xda 0x5a)\n\nHere is what I have tried extracting with:\n\nLinux - ArchLinux - 7zip \nLinux - ArchLinux - Engrampa Archive Manager\nWindows 7 - Winrar (Also on a separate computer)\n\nIn Linux I do get a successful extraction, but then the images are corrupted in some way as described further above. Changing the filename to \".png\" does not fix this.\n\nIn Windows the extraction simply fails, giving a \"CRC failed in -imagepath-. The file is corrupt\" error message.\n\nI haven't found any similar comments on here, which i find strange since it seems to be a computer independent issue. Help would be appreciated, need me some dataz.",
    "183525": "It sounds like your downloaded archive is corrupt, did you download directly? It might have failed early. \n\nTry downloading via the torrent, each chuck is validated as it's downloaded.",
    "183537": "Is there a link to the torrent? I may be missing something. And yes, I only tried the direct download.",
    "183543": "Yeah it's the 3rd file down on the data page. Download the .torrent file then open it in something like [uTorrent][1] to download the data.\n\nI can't see any checksum values on the download page, but the MD5 hash of my (working) copy is: 09630241b306c6fb2093133bcb1319a8\n\n\n  [1]: http://www.utorrent.com/",
    "183545": "Wow, I had totally ignored that. Thanks for pointing it out. I will go ahead and try the torrent and see what happens.",
    "183933": "It still seems to be happening with the torrent on my Linux system. Any thoughts?"
  },
  "source": "meta"
}