{
  "id": 21301,
  "title": "Amazon AWS",
  "url": "/competitions/expedia-hotel-recommendations/discussion/21301",
  "author_name": "",
  "post_date": "2016-05-29T19:22:10.847Z",
  "votes": 3,
  "comment_count": 13,
  "views": 2225,
  "content": "<p>What kind of EC2 did you use to handle this dataset? How about the price and speed? Thank you.</p>",
  "messages": [
    {
      "id": "121790",
      "postDate": "05/29/2016 19:22:10",
      "content": "<p>What kind of EC2 did you use to handle this dataset? How about the price and speed? Thank you.</p>",
      "rawMarkdown": "What kind of EC2 did you use to handle this dataset? How about the price and speed? Thank you.",
      "votes": null
    },
    {
      "id": "121813",
      "postDate": "05/30/2016 03:49:01",
      "content": "<p>@FengLi</p>\n\n<p>I have tried the followings:</p>\n\n<p>m4.4xlarge (16 cpu 64G RAM)</p>\n\n<p>m4.8xlarge(40cpu 160G RAM )</p>\n\n<p>r3.4xlarge (12 cpu 122G RAM)</p>\n\n<p>You can use spot request. It will be far cheaper than the reserved one. Since it is spot request, the price is floating. The price is around ( you can check the price history when you launch spot instance)</p>\n\n<p>0.13~0.25$/hr for m4.4xlarge</p>\n\n<p>0.2-0.4$/hr for m4.8xlarge</p>\n\n<p>0.2-0.4$/hr for r3.4xlarge </p>",
      "rawMarkdown": "FengLi\r\n\r\nI have tried the followings:\r\n\r\nm4.4xlarge (16 cpu 64G RAM)\r\n\r\nm4.8xlarge(40cpu 160G RAM )\r\n\r\nr3.4xlarge (12 cpu 122G RAM)\r\n\r\nYou can use spot request. It will be far cheaper than the reserved one. Since it is spot request, the price is floating. The price is around ( you can check the price history when you launch spot instance)\r\n\r\n0.13~0.25$/hr for m4.4xlarge\r\n\r\n0.2-0.4$/hr for m4.8xlarge\r\n\r\n0.2-0.4$/hr for r3.4xlarge",
      "votes": null
    },
    {
      "id": "122270",
      "postDate": "06/02/2016 15:52:48",
      "content": "<p>@kuan chen How does spot request work? thanks.</p>",
      "rawMarkdown": "kuan chen How does spot request work? thanks.",
      "votes": null
    },
    {
      "id": "122322",
      "postDate": "06/03/2016 03:30:09",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "122323",
      "postDate": "06/03/2016 03:30:12",
      "content": "<p>@zyazzy</p>\n\n<p>In short, spot request is like bidding the price of  available instance in the AWS. So the price is floating because it is determined by the current demand and available resource. To launch AWS spot instance, you can specify it in the AWS EC2 control panel. More details please refer to this <a href=\"http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/spot-requests.html\">document</a>  </p>",
      "rawMarkdown": "zyazzy\r\n\r\nIn short, spot request is like bidding the price of  available instance in the AWS. So the price is floating because it is determined by the current demand and available resource. To launch AWS spot instance, you can specify it in the AWS EC2 control panel. More details please refer to this [document][1]  \r\n\r\n\r\n  [1]: http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/spot-requests.html",
      "votes": null
    },
    {
      "id": "122375",
      "postDate": "06/03/2016 14:41:49",
      "content": "<p>@kuan chen</p>\n\n<p>Thanks. How does it affect our kaggle computations? I normally compute an algorithm for 6-8 hours straight. I understand that spot instances might be interrupted if your bid price is lower than market price, how's your experience so far working with spot instances?</p>",
      "rawMarkdown": "kuan chen\r\n\r\nThanks. How does it affect our kaggle computations? I normally compute an algorithm for 6-8 hours straight. I understand that spot instances might be interrupted if your bid price is lower than market price, how's your experience so far working with spot instances?",
      "votes": null
    },
    {
      "id": "122384",
      "postDate": "06/03/2016 16:04:45",
      "content": "<p>You can check the price trends over the last couple of days and judge roughly where you feel comfortable paying up to. Most times you dont get sharp increases - if your just depending on having it for six hours or so. </p>",
      "rawMarkdown": "You can check the price trends over the last couple of days and judge roughly where you feel comfortable paying up to. Most times you dont get sharp increases - if your just depending on having it for six hours or so.",
      "votes": null
    },
    {
      "id": "122540",
      "postDate": "06/04/2016 22:59:27",
      "content": "<p>@FengLi, @zyazzy\nI exclusively use AWS as my desktop is a Chromebox. My approach is to use Spot Request instance,  mount an EBS volume that has all my data and use AWS DNS server to connect a fully qualified host name to the IP address of the instance. I spend between $10-$20 per month when actively working on a competition.</p>\n\n<p>The steps are: </p>\n\n<ol>\n<li>Pick a region, I chose Sydney as its cheap and close, but cheap is more important.</li>\n<li><p>Use pricing history to get a feel for the spot prices for the class of machine you want to use. </p>\n\n<ul><li>Check which Availability Zone (AZ) in your region is cheapest and most consistently priced for the class of machine you want. For me is ap-southeast-2c</li></ul></li>\n<li><p>I found the General Purpose instances (m4) seems to give the right balance of cores to memory. </p>\n\n<ul><li>Generally I use a m3xlarge, 4 vCPUs and 16GB memory. This typically costs  3.8c and since ap-southeast-2c is rarely used for these servers, the price does not change much. I could big a little more, say 4.5c and ensure it never ever gets bumped</li>\n<li>When I want more performance I use a larger machine. e.g. a   m4.10xlarge has 40 vCPUs and 160GB memory </li></ul></li>\n<li><p>I use the new Spot Request UI  as it's simpler for manually creating a new instance each time. I create this with a standard O/S image ami-0c95b86f </p></li>\n</ol>\n\n<p>5  I use the &quot;user data&quot; field to run a script to automatically configure the machine to:</p>\n\n<ul>\n<li>Mount an EBS volume that has my data and my tools (Anaconda, Python, GraphLab). This way I always have the same apps and data for each new spot instance.</li>\n<li>Update the AWS DNS server to point my the FQDN to the IP address of this host. I do this rather than use elastic IP addresses as they are quite expensive compared to the server. </li>\n</ul>\n\n<p>6 I get a cup of coffee (it takes about 1 min to boot and get ready), then open my Jupyter notebook and a SSH session to the FQDN </p>\n\n<p>It sounds more complex than it is. Once you understand what you are doing and have the EBS volume and the scripts working, it take about 30 sec to request the instance and about 30 sec for it to boot.  Then I've got a clean fast machine that I can take from 1 vCP to 40 vCPUs in about 1 min, i.e. terminate and boot new instance. </p>\n\n<hr>\n\n<p>Here is my script</p>\n\n<pre><code>#!/bin/bash -x\nexec &gt; /tmp/initalise.log  2&gt;&amp;1\n\ncd /tmp\n#Set up AWS IDs so aws cli commands work\nexport AWS_ACCESS_KEY_ID=&lt;your ID&gt;\nexport AWS_SECRET_ACCESS_KEY=&lt;your key&gt;\nexport AWS_DEFAULT_REGION=ap-southeast-2\n\n#Download and run the initialisation script \nwget https://raw.githubusercontent.com/Kevin-McIsaac/bin/master/initalise\nchmod +x initalise\n./initalise vol-9337c448\n</code></pre>",
      "rawMarkdown": "FengLi, @zyazzy\r\nI exclusively use AWS as my desktop is a Chromebox. My approach is to use Spot Request instance,  mount an EBS volume that has all my data and use AWS DNS server to connect a fully qualified host name to the IP address of the instance. I spend between $10-$20 per month when actively working on a competition.\r\n\r\nThe steps are: \r\n \r\n1. Pick a region, I chose Sydney as its cheap and close, but cheap is more important.\r\n2. Use pricing history to get a feel for the spot prices for the class of machine you want to use. \r\n\r\n* Check which Availability Zone (AZ) in your region is cheapest and most consistently priced for the class of machine you want. For me is ap-southeast-2c\r\n\r\n3. I found the General Purpose instances (m4) seems to give the right balance of cores to memory. \r\n\r\n* Generally I use a m3xlarge, 4 vCPUs and 16GB memory. This typically costs  3.8c and since ap-southeast-2c is rarely used for these servers, the price does not change much. I could big a little more, say 4.5c and ensure it never ever gets bumped\r\n* When I want more performance I use a larger machine. e.g. a   m4.10xlarge has 40 vCPUs and 160GB memory \r\n\r\n3.  I use the new Spot Request UI  as it's simpler for manually creating a new instance each time. I create this with a standard O/S image ami-0c95b86f \r\n\r\n5  I use the \"user data\" field to run a script to automatically configure the machine to:\r\n\r\n*  Mount an EBS volume that has my data and my tools (Anaconda, Python, GraphLab). This way I always have the same apps and data for each new spot instance.\r\n* Update the AWS DNS server to point my the FQDN to the IP address of this host. I do this rather than use elastic IP addresses as they are quite expensive compared to the server. \r\n\r\n6 I get a cup of coffee (it takes about 1 min to boot and get ready), then open my Jupyter notebook and a SSH session to the FQDN \r\n\r\nIt sounds more complex than it is. Once you understand what you are doing and have the EBS volume and the scripts working, it take about 30 sec to request the instance and about 30 sec for it to boot.  Then I've got a clean fast machine that I can take from 1 vCP to 40 vCPUs in about 1 min, i.e. terminate and boot new instance. \r\n\r\n---------------\r\nHere is my script\r\n\r\n    #!/bin/bash -x\r\n    exec > /tmp/initalise.log  2>&1\r\n    \r\n    cd /tmp\r\n    #Set up AWS IDs so aws cli commands work\r\n    export AWS_ACCESS_KEY_ID=<your ID>\r\n    export AWS_SECRET_ACCESS_KEY=<your key>\r\n    export AWS_DEFAULT_REGION=ap-southeast-2\r\n    \r\n    #Download and run the initialisation script \r\n    wget https://raw.githubusercontent.com/Kevin-McIsaac/bin/master/initalise\r\n    chmod +x initalise\r\n    ./initalise vol-9337c448",
      "votes": null
    },
    {
      "id": "122715",
      "postDate": "06/06/2016 19:00:00",
      "content": "<p>Hello, \nI signed up and followed the steps to install anaconda (followed <a href=\"http://www.grant-mckinnon.com/?p=6\">http://www.grant-mckinnon.com/?p=6</a>) and ran scripts. Since free-tier does not accomodate the data sizes of the competition, I decided to rent a more powerful instance. Do I have to re-install anaconda and send the files again, or can I just move the volumes (created with the free-tier) to this new instance? \nThanks!</p>",
      "rawMarkdown": "Hello, \r\nI signed up and followed the steps to install anaconda (followed http://www.grant-mckinnon.com/?p=6) and ran scripts. Since free-tier does not accomodate the data sizes of the competition, I decided to rent a more powerful instance. Do I have to re-install anaconda and send the files again, or can I just move the volumes (created with the free-tier) to this new instance? \r\nThanks!",
      "votes": null
    },
    {
      "id": "122752",
      "postDate": "06/06/2016 23:34:40",
      "content": "<p>The simplest way to do this, is select your instance, ensure it is stopped, then use Actions&gt; Instance Settings&gt;Change Instance Type  to change this to a large machine. You then start the instance and the sever now has the new hardware.</p>",
      "rawMarkdown": "The simplest way to do this, is select your instance, ensure it is stopped, then use Actions> Instance Settings>Change Instance Type  to change this to a large machine. You then start the instance and the sever now has the new hardware.",
      "votes": null
    },
    {
      "id": "122781",
      "postDate": "06/07/2016 04:50:42",
      "content": "<p>@Fractal Feelings, thanks! I had chosen c3.2xlarge with 15G RAM, and getting Memory Error when reading csv into pandas. Not sure why this happens. In my MacOS, the train.csv is loaded with no problem. </p>",
      "rawMarkdown": "Fractal Feelings, thanks! I had chosen c3.2xlarge with 15G RAM, and getting Memory Error when reading csv into pandas. Not sure why this happens. In my MacOS, the train.csv is loaded with no problem.",
      "votes": null
    },
    {
      "id": "122876",
      "postDate": "06/07/2016 23:42:10",
      "content": "<p>BTW if I'm not using the free tier, I use a general purpose (m4) class machine as I find this the best price performacne</p>",
      "rawMarkdown": "BTW if I'm not using the free tier, I use a general purpose (m4) class machine as I find this the best price performacne",
      "votes": null
    },
    {
      "id": "122972",
      "postDate": "06/08/2016 22:59:56",
      "content": "<p>I SET UP MY FIRST AWS INSTANCE AND REMOTED INTO A PYTHON NOTEBOOK SERVER TODAY!!!!!!!!!!</p>\n\n<p>Woooooooooohoooooooooo! Now I get to start kaggling! Lame internet speed and processing no more.</p>\n\n<p>I even did it in ubuntu intead of windows (what I'm used to) because I figured I'm going to have learn unix at some point. YEEEEEEEEESSSSS!!!!!!!</p>",
      "rawMarkdown": "I SET UP MY FIRST AWS INSTANCE AND REMOTED INTO A PYTHON NOTEBOOK SERVER TODAY!!!!!!!!!!\r\n\r\nWoooooooooohoooooooooo! Now I get to start kaggling! Lame internet speed and processing no more.\r\n\r\nI even did it in ubuntu intead of windows (what I'm used to) because I figured I'm going to have learn unix at some point. YEEEEEEEEESSSSS!!!!!!!",
      "votes": null
    },
    {
      "id": "360354",
      "postDate": "07/22/2018 08:24:48",
      "content": "<p>How did you get the Kaggle API Key onto the AWS instance to be able to download datasets?</p>",
      "rawMarkdown": "How did you get the Kaggle API Key onto the AWS instance to be able to download datasets?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 121813,
      "author_name": "kuanchen",
      "author_url": "",
      "post_date": "05/30/2016 03:49:01",
      "content": "<p>@FengLi</p>\n\n<p>I have tried the followings:</p>\n\n<p>m4.4xlarge (16 cpu 64G RAM)</p>\n\n<p>m4.8xlarge(40cpu 160G RAM )</p>\n\n<p>r3.4xlarge (12 cpu 122G RAM)</p>\n\n<p>You can use spot request. It will be far cheaper than the reserved one. Since it is spot request, the price is floating. The price is around ( you can check the price history when you launch spot instance)</p>\n\n<p>0.13~0.25$/hr for m4.4xlarge</p>\n\n<p>0.2-0.4$/hr for m4.8xlarge</p>\n\n<p>0.2-0.4$/hr for r3.4xlarge </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122270,
      "author_name": "zyazzy",
      "author_url": "",
      "post_date": "06/02/2016 15:52:48",
      "content": "<p>@kuan chen How does spot request work? thanks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122322,
      "author_name": "kuanchen",
      "author_url": "",
      "post_date": "06/03/2016 03:30:09",
      "content": "",
      "votes": null,
      "replies": []
    },
    {
      "id": 122323,
      "author_name": "kuanchen",
      "author_url": "",
      "post_date": "06/03/2016 03:30:12",
      "content": "<p>@zyazzy</p>\n\n<p>In short, spot request is like bidding the price of  available instance in the AWS. So the price is floating because it is determined by the current demand and available resource. To launch AWS spot instance, you can specify it in the AWS EC2 control panel. More details please refer to this <a href=\"http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/spot-requests.html\">document</a>  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122375,
      "author_name": "zyazzy",
      "author_url": "",
      "post_date": "06/03/2016 14:41:49",
      "content": "<p>@kuan chen</p>\n\n<p>Thanks. How does it affect our kaggle computations? I normally compute an algorithm for 6-8 hours straight. I understand that spot instances might be interrupted if your bid price is lower than market price, how's your experience so far working with spot instances?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122384,
      "author_name": "darraghdog",
      "author_url": "",
      "post_date": "06/03/2016 16:04:45",
      "content": "<p>You can check the price trends over the last couple of days and judge roughly where you feel comfortable paying up to. Most times you dont get sharp increases - if your just depending on having it for six hours or so. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122540,
      "author_name": "kevinmcisaac",
      "author_url": "",
      "post_date": "06/04/2016 22:59:27",
      "content": "<p>@FengLi, @zyazzy\nI exclusively use AWS as my desktop is a Chromebox. My approach is to use Spot Request instance,  mount an EBS volume that has all my data and use AWS DNS server to connect a fully qualified host name to the IP address of the instance. I spend between $10-$20 per month when actively working on a competition.</p>\n\n<p>The steps are: </p>\n\n<ol>\n<li>Pick a region, I chose Sydney as its cheap and close, but cheap is more important.</li>\n<li><p>Use pricing history to get a feel for the spot prices for the class of machine you want to use. </p>\n\n<ul><li>Check which Availability Zone (AZ) in your region is cheapest and most consistently priced for the class of machine you want. For me is ap-southeast-2c</li></ul></li>\n<li><p>I found the General Purpose instances (m4) seems to give the right balance of cores to memory. </p>\n\n<ul><li>Generally I use a m3xlarge, 4 vCPUs and 16GB memory. This typically costs  3.8c and since ap-southeast-2c is rarely used for these servers, the price does not change much. I could big a little more, say 4.5c and ensure it never ever gets bumped</li>\n<li>When I want more performance I use a larger machine. e.g. a   m4.10xlarge has 40 vCPUs and 160GB memory </li></ul></li>\n<li><p>I use the new Spot Request UI  as it's simpler for manually creating a new instance each time. I create this with a standard O/S image ami-0c95b86f </p></li>\n</ol>\n\n<p>5  I use the &quot;user data&quot; field to run a script to automatically configure the machine to:</p>\n\n<ul>\n<li>Mount an EBS volume that has my data and my tools (Anaconda, Python, GraphLab). This way I always have the same apps and data for each new spot instance.</li>\n<li>Update the AWS DNS server to point my the FQDN to the IP address of this host. I do this rather than use elastic IP addresses as they are quite expensive compared to the server. </li>\n</ul>\n\n<p>6 I get a cup of coffee (it takes about 1 min to boot and get ready), then open my Jupyter notebook and a SSH session to the FQDN </p>\n\n<p>It sounds more complex than it is. Once you understand what you are doing and have the EBS volume and the scripts working, it take about 30 sec to request the instance and about 30 sec for it to boot.  Then I've got a clean fast machine that I can take from 1 vCP to 40 vCPUs in about 1 min, i.e. terminate and boot new instance. </p>\n\n<hr>\n\n<p>Here is my script</p>\n\n<pre><code>#!/bin/bash -x\nexec &gt; /tmp/initalise.log  2&gt;&amp;1\n\ncd /tmp\n#Set up AWS IDs so aws cli commands work\nexport AWS_ACCESS_KEY_ID=&lt;your ID&gt;\nexport AWS_SECRET_ACCESS_KEY=&lt;your key&gt;\nexport AWS_DEFAULT_REGION=ap-southeast-2\n\n#Download and run the initialisation script \nwget https://raw.githubusercontent.com/Kevin-McIsaac/bin/master/initalise\nchmod +x initalise\n./initalise vol-9337c448\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122715,
      "author_name": "hitoshinagano",
      "author_url": "",
      "post_date": "06/06/2016 19:00:00",
      "content": "<p>Hello, \nI signed up and followed the steps to install anaconda (followed <a href=\"http://www.grant-mckinnon.com/?p=6\">http://www.grant-mckinnon.com/?p=6</a>) and ran scripts. Since free-tier does not accomodate the data sizes of the competition, I decided to rent a more powerful instance. Do I have to re-install anaconda and send the files again, or can I just move the volumes (created with the free-tier) to this new instance? \nThanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122752,
      "author_name": "kevinmcisaac",
      "author_url": "",
      "post_date": "06/06/2016 23:34:40",
      "content": "<p>The simplest way to do this, is select your instance, ensure it is stopped, then use Actions&gt; Instance Settings&gt;Change Instance Type  to change this to a large machine. You then start the instance and the sever now has the new hardware.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122781,
      "author_name": "hitoshinagano",
      "author_url": "",
      "post_date": "06/07/2016 04:50:42",
      "content": "<p>@Fractal Feelings, thanks! I had chosen c3.2xlarge with 15G RAM, and getting Memory Error when reading csv into pandas. Not sure why this happens. In my MacOS, the train.csv is loaded with no problem. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122876,
      "author_name": "kevinmcisaac",
      "author_url": "",
      "post_date": "06/07/2016 23:42:10",
      "content": "<p>BTW if I'm not using the free tier, I use a general purpose (m4) class machine as I find this the best price performacne</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122972,
      "author_name": "lit041000",
      "author_url": "",
      "post_date": "06/08/2016 22:59:56",
      "content": "<p>I SET UP MY FIRST AWS INSTANCE AND REMOTED INTO A PYTHON NOTEBOOK SERVER TODAY!!!!!!!!!!</p>\n\n<p>Woooooooooohoooooooooo! Now I get to start kaggling! Lame internet speed and processing no more.</p>\n\n<p>I even did it in ubuntu intead of windows (what I'm used to) because I figured I'm going to have learn unix at some point. YEEEEEEEEESSSSS!!!!!!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 360354,
      "author_name": "saadk408",
      "author_url": "",
      "post_date": "07/22/2018 08:24:48",
      "content": "<p>How did you get the Kaggle API Key onto the AWS instance to be able to download datasets?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "121790": "What kind of EC2 did you use to handle this dataset? How about the price and speed? Thank you.",
    "121813": "FengLi\r\n\r\nI have tried the followings:\r\n\r\nm4.4xlarge (16 cpu 64G RAM)\r\n\r\nm4.8xlarge(40cpu 160G RAM )\r\n\r\nr3.4xlarge (12 cpu 122G RAM)\r\n\r\nYou can use spot request. It will be far cheaper than the reserved one. Since it is spot request, the price is floating. The price is around ( you can check the price history when you launch spot instance)\r\n\r\n0.13~0.25$/hr for m4.4xlarge\r\n\r\n0.2-0.4$/hr for m4.8xlarge\r\n\r\n0.2-0.4$/hr for r3.4xlarge",
    "122270": "kuan chen How does spot request work? thanks.",
    "122322": "",
    "122323": "zyazzy\r\n\r\nIn short, spot request is like bidding the price of  available instance in the AWS. So the price is floating because it is determined by the current demand and available resource. To launch AWS spot instance, you can specify it in the AWS EC2 control panel. More details please refer to this [document][1]  \r\n\r\n\r\n  [1]: http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/spot-requests.html",
    "122375": "kuan chen\r\n\r\nThanks. How does it affect our kaggle computations? I normally compute an algorithm for 6-8 hours straight. I understand that spot instances might be interrupted if your bid price is lower than market price, how's your experience so far working with spot instances?",
    "122384": "You can check the price trends over the last couple of days and judge roughly where you feel comfortable paying up to. Most times you dont get sharp increases - if your just depending on having it for six hours or so.",
    "122540": "FengLi, @zyazzy\r\nI exclusively use AWS as my desktop is a Chromebox. My approach is to use Spot Request instance,  mount an EBS volume that has all my data and use AWS DNS server to connect a fully qualified host name to the IP address of the instance. I spend between $10-$20 per month when actively working on a competition.\r\n\r\nThe steps are: \r\n \r\n1. Pick a region, I chose Sydney as its cheap and close, but cheap is more important.\r\n2. Use pricing history to get a feel for the spot prices for the class of machine you want to use. \r\n\r\n* Check which Availability Zone (AZ) in your region is cheapest and most consistently priced for the class of machine you want. For me is ap-southeast-2c\r\n\r\n3. I found the General Purpose instances (m4) seems to give the right balance of cores to memory. \r\n\r\n* Generally I use a m3xlarge, 4 vCPUs and 16GB memory. This typically costs  3.8c and since ap-southeast-2c is rarely used for these servers, the price does not change much. I could big a little more, say 4.5c and ensure it never ever gets bumped\r\n* When I want more performance I use a larger machine. e.g. a   m4.10xlarge has 40 vCPUs and 160GB memory \r\n\r\n3.  I use the new Spot Request UI  as it's simpler for manually creating a new instance each time. I create this with a standard O/S image ami-0c95b86f \r\n\r\n5  I use the \"user data\" field to run a script to automatically configure the machine to:\r\n\r\n*  Mount an EBS volume that has my data and my tools (Anaconda, Python, GraphLab). This way I always have the same apps and data for each new spot instance.\r\n* Update the AWS DNS server to point my the FQDN to the IP address of this host. I do this rather than use elastic IP addresses as they are quite expensive compared to the server. \r\n\r\n6 I get a cup of coffee (it takes about 1 min to boot and get ready), then open my Jupyter notebook and a SSH session to the FQDN \r\n\r\nIt sounds more complex than it is. Once you understand what you are doing and have the EBS volume and the scripts working, it take about 30 sec to request the instance and about 30 sec for it to boot.  Then I've got a clean fast machine that I can take from 1 vCP to 40 vCPUs in about 1 min, i.e. terminate and boot new instance. \r\n\r\n---------------\r\nHere is my script\r\n\r\n    #!/bin/bash -x\r\n    exec > /tmp/initalise.log  2>&1\r\n    \r\n    cd /tmp\r\n    #Set up AWS IDs so aws cli commands work\r\n    export AWS_ACCESS_KEY_ID=<your ID>\r\n    export AWS_SECRET_ACCESS_KEY=<your key>\r\n    export AWS_DEFAULT_REGION=ap-southeast-2\r\n    \r\n    #Download and run the initialisation script \r\n    wget https://raw.githubusercontent.com/Kevin-McIsaac/bin/master/initalise\r\n    chmod +x initalise\r\n    ./initalise vol-9337c448",
    "122715": "Hello, \r\nI signed up and followed the steps to install anaconda (followed http://www.grant-mckinnon.com/?p=6) and ran scripts. Since free-tier does not accomodate the data sizes of the competition, I decided to rent a more powerful instance. Do I have to re-install anaconda and send the files again, or can I just move the volumes (created with the free-tier) to this new instance? \r\nThanks!",
    "122752": "The simplest way to do this, is select your instance, ensure it is stopped, then use Actions> Instance Settings>Change Instance Type  to change this to a large machine. You then start the instance and the sever now has the new hardware.",
    "122781": "Fractal Feelings, thanks! I had chosen c3.2xlarge with 15G RAM, and getting Memory Error when reading csv into pandas. Not sure why this happens. In my MacOS, the train.csv is loaded with no problem.",
    "122876": "BTW if I'm not using the free tier, I use a general purpose (m4) class machine as I find this the best price performacne",
    "122972": "I SET UP MY FIRST AWS INSTANCE AND REMOTED INTO A PYTHON NOTEBOOK SERVER TODAY!!!!!!!!!!\r\n\r\nWoooooooooohoooooooooo! Now I get to start kaggling! Lame internet speed and processing no more.\r\n\r\nI even did it in ubuntu intead of windows (what I'm used to) because I figured I'm going to have learn unix at some point. YEEEEEEEEESSSSS!!!!!!!",
    "360354": "How did you get the Kaggle API Key onto the AWS instance to be able to download datasets?"
  },
  "source": "meta"
}