{
  "id": 71950,
  "title": "Regarding fullsize image data",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/71950",
  "author_name": "",
  "post_date": "2018-11-18T17:29:07.307095800Z",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>The size of the images fed into the neural net is quite significant in this competition so I expect a lot of people will be trying out larger image sizes.  This means using the fullsize data.  These images are stored in 7z archives and take up a lot of disk space when decompressed.  Not only do you need enough space for the decompressed files, you also need space for the archive while decompressing.</p>\n\n<p>To avoid decompressing the archive I was thinking that one could use some library to extract individual files from it, thereby avoiding having to decompress the entire archive at once.  Unfortunately I couldn't find a python package that would allow me to do it.  The two packages that I found either had some problem during installation or during importing.</p>\n\n<p>I think (not absolutely sure) there are mature zip libraries available for python that would allow this type of manipulation on zip archives.  If that is the case then it would be great if the organizers could upload zip archives for the contestants to use as well.  This could also be considered for future competitions where archives with large amounts of data are being used.</p>",
  "messages": [
    {
      "id": "423609",
      "postDate": "11/18/2018 17:29:07",
      "content": "<p>The size of the images fed into the neural net is quite significant in this competition so I expect a lot of people will be trying out larger image sizes.  This means using the fullsize data.  These images are stored in 7z archives and take up a lot of disk space when decompressed.  Not only do you need enough space for the decompressed files, you also need space for the archive while decompressing.</p>\n\n<p>To avoid decompressing the archive I was thinking that one could use some library to extract individual files from it, thereby avoiding having to decompress the entire archive at once.  Unfortunately I couldn't find a python package that would allow me to do it.  The two packages that I found either had some problem during installation or during importing.</p>\n\n<p>I think (not absolutely sure) there are mature zip libraries available for python that would allow this type of manipulation on zip archives.  If that is the case then it would be great if the organizers could upload zip archives for the contestants to use as well.  This could also be considered for future competitions where archives with large amounts of data are being used.</p>",
      "rawMarkdown": "The size of the images fed into the neural net is quite significant in this competition so I expect a lot of people will be trying out larger image sizes.  This means using the fullsize data.  These images are stored in 7z archives and take up a lot of disk space when decompressed.  Not only do you need enough space for the decompressed files, you also need space for the archive while decompressing.\n\nTo avoid decompressing the archive I was thinking that one could use some library to extract individual files from it, thereby avoiding having to decompress the entire archive at once.  Unfortunately I couldn't find a python package that would allow me to do it.  The two packages that I found either had some problem during installation or during importing.\n\nI think (not absolutely sure) there are mature zip libraries available for python that would allow this type of manipulation on zip archives.  If that is the case then it would be great if the organizers could upload zip archives for the contestants to use as well.  This could also be considered for future competitions where archives with large amounts of data are being used.",
      "votes": null
    },
    {
      "id": "423623",
      "postDate": "11/18/2018 18:23:15",
      "content": "<p>If your idea is to run tbe training directly on the zipped content, unfortunately this is not a great idea. The amount of io and cpu involved is prohibiting. If you are running the training locally, get a 99$ 2tb external usb 3 drive and archive the zipped archives there. Its a good investment anyway for backups... </p>",
      "rawMarkdown": "If your idea is to run tbe training directly on the zipped content, unfortunately this is not a great idea. The amount of io and cpu involved is prohibiting. If you are running the training locally, get a 99$ 2tb external usb 3 drive and archive the zipped archives there. Its a good investment anyway for backups...",
      "votes": null
    },
    {
      "id": "423680",
      "postDate": "11/18/2018 21:42:08",
      "content": "<p>No, I wasn't planning on training directly on the zipped content.  It's to make it easier to access the files without having to unzip everything on the hard drive and take up a lot of space, as explained in my original post.</p>",
      "rawMarkdown": "No, I wasn't planning on training directly on the zipped content.  It's to make it easier to access the files without having to unzip everything on the hard drive and take up a lot of space, as explained in my original post.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 423623,
      "author_name": "moshel",
      "author_url": "",
      "post_date": "11/18/2018 18:23:15",
      "content": "<p>If your idea is to run tbe training directly on the zipped content, unfortunately this is not a great idea. The amount of io and cpu involved is prohibiting. If you are running the training locally, get a 99$ 2tb external usb 3 drive and archive the zipped archives there. Its a good investment anyway for backups... </p>",
      "votes": null,
      "replies": [
        {
          "id": 423680,
          "author_name": "robertkag",
          "author_url": "",
          "post_date": "11/18/2018 21:42:08",
          "content": "<p>No, I wasn't planning on training directly on the zipped content.  It's to make it easier to access the files without having to unzip everything on the hard drive and take up a lot of space, as explained in my original post.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "423609": "The size of the images fed into the neural net is quite significant in this competition so I expect a lot of people will be trying out larger image sizes.  This means using the fullsize data.  These images are stored in 7z archives and take up a lot of disk space when decompressed.  Not only do you need enough space for the decompressed files, you also need space for the archive while decompressing.\n\nTo avoid decompressing the archive I was thinking that one could use some library to extract individual files from it, thereby avoiding having to decompress the entire archive at once.  Unfortunately I couldn't find a python package that would allow me to do it.  The two packages that I found either had some problem during installation or during importing.\n\nI think (not absolutely sure) there are mature zip libraries available for python that would allow this type of manipulation on zip archives.  If that is the case then it would be great if the organizers could upload zip archives for the contestants to use as well.  This could also be considered for future competitions where archives with large amounts of data are being used.",
    "423623": "If your idea is to run tbe training directly on the zipped content, unfortunately this is not a great idea. The amount of io and cpu involved is prohibiting. If you are running the training locally, get a 99$ 2tb external usb 3 drive and archive the zipped archives there. Its a good investment anyway for backups...",
    "423680": "No, I wasn't planning on training directly on the zipped content.  It's to make it easier to access the files without having to unzip everything on the hard drive and take up a lot of space, as explained in my original post."
  },
  "source": "meta"
}