{
  "id": 10224,
  "title": "where to save and extract these big zipped file?",
  "url": "/competitions/seizure-prediction/discussion/10224",
  "author_name": "",
  "post_date": "2014-09-04T22:39:57.430Z",
  "votes": null,
  "comment_count": 2,
  "views": 1118,
  "content": "<p>I am new to Kaggle competitions.</p>\n\n<p>One question, where do you save and extract these big zipped data files?</p>\n<p>Do you all handle them in servers?</p>\n<p>They can hardly be handled on pc hard disk.</p>",
  "messages": [
    {
      "id": "53126",
      "postDate": "09/04/2014 22:39:57",
      "content": "<p>I am new to Kaggle competitions.</p>\n\n<p>One question, where do you save and extract these big zipped data files?</p>\n<p>Do you all handle them in servers?</p>\n<p>They can hardly be handled on pc hard disk.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "53128",
      "postDate": "09/04/2014 22:59:48",
      "content": "<p>bigger boat?</p>\n<p>or run everything on Amazon AWS and keep everything on a Volume</p>\n<p>or&nbsp;can save the files on S3 and re-load from S3 to the SSD each time you start the machine.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "53132",
      "postDate": "09/04/2014 23:29:52",
      "content": "<p>This way is definitely not &quot;optimal&quot;, but how I tackle it.</p>\n<p>I keep the zipped files on a local network drive. This means I can redownload with a much higher speed, if I need to re-engineer features.</p>\n<p>I try to generate as many features as I can think of, based on CV of 1 or 2 patients (or read about in the last&nbsp;competition thread) and store this as per-patient train and test sets for every patient, switching out/deleting the source data, till only reduced datasets are left (which should fit on SSD HD).</p>\n<p>Then feature selection, modeling etc. on these datasets. Submitting, receiving a terrible score due to a mistake, giving up hope, before retrying everything :).</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 53128,
      "author_name": "udibr1",
      "author_url": "",
      "post_date": "09/04/2014 22:59:48",
      "content": "<p>bigger boat?</p>\n<p>or run everything on Amazon AWS and keep everything on a Volume</p>\n<p>or&nbsp;can save the files on S3 and re-load from S3 to the SSD each time you start the machine.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 53132,
      "author_name": "triskelion",
      "author_url": "",
      "post_date": "09/04/2014 23:29:52",
      "content": "<p>This way is definitely not &quot;optimal&quot;, but how I tackle it.</p>\n<p>I keep the zipped files on a local network drive. This means I can redownload with a much higher speed, if I need to re-engineer features.</p>\n<p>I try to generate as many features as I can think of, based on CV of 1 or 2 patients (or read about in the last&nbsp;competition thread) and store this as per-patient train and test sets for every patient, switching out/deleting the source data, till only reduced datasets are left (which should fit on SSD HD).</p>\n<p>Then feature selection, modeling etc. on these datasets. Submitting, receiving a terrible score due to a mistake, giving up hope, before retrying everything :).</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "53126": "",
    "53128": "",
    "53132": ""
  },
  "source": "meta"
}