{
  "id": 20509,
  "title": "Uploading files to AWS - R studio",
  "url": "/competitions/expedia-hotel-recommendations/discussion/20509",
  "author_name": "",
  "post_date": "2016-04-28T13:50:14.980Z",
  "votes": 3,
  "comment_count": 15,
  "views": 2271,
  "content": "<p>Hello,</p>\n\n<p>I have set up R-Studio on AWS and have been able to upload destinations.csv file, using the option to upload file in R studio, it uploaded fairly quickly. However train.csv upload is taking forever.</p>\n\n<p>Can you please let me know, if there is a quicker way to upload files.</p>\n\n<p>Thanks,</p>",
  "messages": [
    {
      "id": "117318",
      "postDate": "04/28/2016 13:50:14",
      "content": "<p>Hello,</p>\n\n<p>I have set up R-Studio on AWS and have been able to upload destinations.csv file, using the option to upload file in R studio, it uploaded fairly quickly. However train.csv upload is taking forever.</p>\n\n<p>Can you please let me know, if there is a quicker way to upload files.</p>\n\n<p>Thanks,</p>",
      "rawMarkdown": "Hello,\r\n\r\nI have set up R-Studio on AWS and have been able to upload destinations.csv file, using the option to upload file in R studio, it uploaded fairly quickly. However train.csv upload is taking forever.\r\n\r\nCan you please let me know, if there is a quicker way to upload files.\r\n\r\nThanks,",
      "votes": null
    },
    {
      "id": "117327",
      "postDate": "04/28/2016 15:21:53",
      "content": "<p>It's a little bit tricky, but you can open the Network tab on Chrome developer tools, click to download your train file (on Kaggle website) and watch for the url it uses. Than stop the download, right click the request on the Network tab and choose &quot;Copy as cUrl&quot;. This copies the exact command you can use on a remote machine to download the file (just remember to pipe the output the a file).</p>\n\n<p>I used this method to get files on a cloud computer.</p>",
      "rawMarkdown": "It's a little bit tricky, but you can open the Network tab on Chrome developer tools, click to download your train file (on Kaggle website) and watch for the url it uses. Than stop the download, right click the request on the Network tab and choose \"Copy as cUrl\". This copies the exact command you can use on a remote machine to download the file (just remember to pipe the output the a file).\r\n\r\nI used this method to get files on a cloud computer.",
      "votes": null
    },
    {
      "id": "117344",
      "postDate": "04/28/2016 16:17:13",
      "content": "<p>Thanks Bruno. Let me try .</p>",
      "rawMarkdown": "Thanks Bruno. Let me try .",
      "votes": null
    },
    {
      "id": "117352",
      "postDate": "04/28/2016 17:12:34",
      "content": "<p>Have you seen this?</p>\n\n<p><a href=\"http://www.louisaslett.com/RStudio_AMI/\">http://www.louisaslett.com/RStudio_AMI/</a></p>\n\n<p>I owe Louis money for the time he's saved me.</p>\n\n<p>I start up one of these instances, and while it's installing xgboost etc, I scp the files I need over.  It's quick and easy.</p>",
      "rawMarkdown": "Have you seen this?\r\n\r\nhttp://www.louisaslett.com/RStudio_AMI/\r\n\r\nI owe Louis money for the time he's saved me.\r\n\r\nI start up one of these instances, and while it's installing xgboost etc, I scp the files I need over.  It's quick and easy.",
      "votes": null
    },
    {
      "id": "117491",
      "postDate": "04/29/2016 11:25:40",
      "content": "<p>Thanks wally. yes, I did see.</p>\n\n<p>I have uploaded using filezilla. It was easy as well. </p>\n\n<p><a href=\"http://angus.readthedocs.io/en/2014/amazon/transfer-files-between-instance.html\">http://angus.readthedocs.io/en/2014/amazon/transfer-files-between-instance.html</a></p>\n\n<p>when I read test set, it has only 1,875,015 rows, is it not 2,528,243  rows are expected ? </p>",
      "rawMarkdown": "Thanks wally. yes, I did see.\r\n\r\nI have uploaded using filezilla. It was easy as well. \r\n\r\nhttp://angus.readthedocs.io/en/2014/amazon/transfer-files-between-instance.html\r\n\r\nwhen I read test set, it has only 1,875,015 rows, is it not 2,528,243  rows are expected ?",
      "votes": null
    },
    {
      "id": "117498",
      "postDate": "04/29/2016 12:43:38",
      "content": "<p>Yes, 2.5... is what I have.</p>",
      "rawMarkdown": "Yes, 2.5... is what I have.",
      "votes": null
    },
    {
      "id": "117613",
      "postDate": "04/29/2016 20:43:19",
      "content": "<p>How is it done if you develop in python? </p>",
      "rawMarkdown": "How is it done if you develop in python?",
      "votes": null
    },
    {
      "id": "117659",
      "postDate": "04/30/2016 02:04:37",
      "content": "<p>read the document @ <a href=\"http://angus.readthedocs.io/en/2014/amazon/transfer-files-between-instance.html\">http://angus.readthedocs.io/en/2014/amazon/transfer-files-between-instance.html</a>, it talks about transferring files from local PC to AWS Instance, it has nothing to do with R or python.</p>",
      "rawMarkdown": "read the document @ http://angus.readthedocs.io/en/2014/amazon/transfer-files-between-instance.html, it talks about transferring files from local PC to AWS Instance, it has nothing to do with R or python.",
      "votes": null
    },
    {
      "id": "117734",
      "postDate": "04/30/2016 16:05:22",
      "content": "<p>@Bruno - excellent recommendation from you. It took less than one minute to download the .5GB file to my instance on AWS (US-EAST). Also:\nexpedia_train &lt;- fread(&quot;gunzip -c train.csv.gz&quot;, header=TRUE)\nRead 37670293 rows and 24 (of 24) columns from 3.791 GB file in 00:02:02</p>\n\n<p>[quote=Bruno G. do Amaral;117327]</p>\n\n<p>It's a little bit tricky, but you can open the Network tab on Chrome developer tools, click to download your train file (on Kaggle website) and watch for the url it uses. Than stop the download, right click the request on the Network tab and choose &quot;Copy as cUrl&quot;. This copies the exact command you can use on a remote machine to download the file (just remember to pipe the output the a file).</p>\n\n<p>I used this method to get files on a cloud computer.</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "Bruno - excellent recommendation from you. It took less than one minute to download the .5GB file to my instance on AWS (US-EAST). Also:\r\nexpedia_train <- fread(\"gunzip -c train.csv.gz\", header=TRUE)\r\nRead 37670293 rows and 24 (of 24) columns from 3.791 GB file in 00:02:02\r\n\r\n[quote=Bruno G. do Amaral;117327]\r\n\r\nIt's a little bit tricky, but you can open the Network tab on Chrome developer tools, click to download your train file (on Kaggle website) and watch for the url it uses. Than stop the download, right click the request on the Network tab and choose \"Copy as cUrl\". This copies the exact command you can use on a remote machine to download the file (just remember to pipe the output the a file).\r\n\r\nI used this method to get files on a cloud computer.\r\n\r\n[/quote]",
      "votes": null
    },
    {
      "id": "118302",
      "postDate": "05/03/2016 07:53:18",
      "content": "<p>Dear Friends\nthis form of reading data does not work on windows system. does it?</p>",
      "rawMarkdown": "Dear Friends\r\nthis form of reading data does not work on windows system. does it?",
      "votes": null
    },
    {
      "id": "118307",
      "postDate": "05/03/2016 08:15:12",
      "content": "<p>@Maryam, I use windows, and I have uploaded using Filezilla, I understand you are asking about the other method, but I did not try that.</p>",
      "rawMarkdown": "Maryam, I use windows, and I have uploaded using Filezilla, I understand you are asking about the other method, but I did not try that.",
      "votes": null
    },
    {
      "id": "119160",
      "postDate": "05/07/2016 17:49:19",
      "content": "<p>If I understood the question correctly, you have downloaded the data on your local machine. And now you were trying to upload it to an AWS instance. Instead why not download it from kaggle directly to the aws instance: </p>\n\n<pre><code>wget --load-cookies=cookies.txt &lt;url to file to download&gt;\n</code></pre>\n\n<p>cookies.txt &gt;&gt; has the cookie data from the kaggle website</p>",
      "rawMarkdown": "If I understood the question correctly, you have downloaded the data on your local machine. And now you were trying to upload it to an AWS instance. Instead why not download it from kaggle directly to the aws instance: \r\n\r\n    wget --load-cookies=cookies.txt <url to file to download>\r\n\r\ncookies.txt >> has the cookie data from the kaggle website",
      "votes": null
    },
    {
      "id": "119203",
      "postDate": "05/08/2016 01:42:41",
      "content": "<p>Thanks, Skylord -  Each file download took less than 1 second!</p>\n\n<p>Edit: maybe because something isn't quite right...files won't unzip :(</p>\n\n<p>Edit2: It was a problem with my cookies file. Fixed now, and train file took a mere 85 seconds.</p>\n\n<p><img src=\"https://dl.dropboxusercontent.com/u/15775348/ffd.PNG\" alt=\"enter image description here\" title></p>",
      "rawMarkdown": "Thanks, Skylord -  Each file download took less than 1 second!\r\n\r\nEdit: maybe because something isn't quite right...files won't unzip :(\r\n\r\nEdit2: It was a problem with my cookies file. Fixed now, and train file took a mere 85 seconds.\r\n\r\n![enter image description here][1]\r\n\r\n\r\n  [1]: https://dl.dropboxusercontent.com/u/15775348/ffd.PNG",
      "votes": null
    },
    {
      "id": "119315",
      "postDate": "05/09/2016 06:24:03",
      "content": "<p>JohnM - This was shared by someone in an earlier competitions. Just repeating it over here. Glad you found it useful. </p>",
      "rawMarkdown": "JohnM - This was shared by someone in an earlier competitions. Just repeating it over here. Glad you found it useful.",
      "votes": null
    },
    {
      "id": "119317",
      "postDate": "05/09/2016 06:54:05",
      "content": "<p>You can simply use a cookie-export extension in Chrome. I use the following:\n<a href=\"https://chrome.google.com/webstore/detail/cookietxt-export/lopabhfecdfhgogdbojmaicoicjekelh\">https://chrome.google.com/webstore/detail/cookietxt-export/lopabhfecdfhgogdbojmaicoicjekelh</a></p>\n\n<p>Then just load the cookies as SkyLord commented above.</p>",
      "rawMarkdown": "You can simply use a cookie-export extension in Chrome. I use the following:\r\nhttps://chrome.google.com/webstore/detail/cookietxt-export/lopabhfecdfhgogdbojmaicoicjekelh\r\n\r\nThen just load the cookies as SkyLord commented above.",
      "votes": null
    },
    {
      "id": "119357",
      "postDate": "05/09/2016 14:03:12",
      "content": "<p>for a  firefox add-in - \n<a href=\"https://addons.mozilla.org/en-US/firefox/addon/export-cookies/\">https://addons.mozilla.org/en-US/firefox/addon/export-cookies/</a></p>",
      "rawMarkdown": "for a  firefox add-in - \r\nhttps://addons.mozilla.org/en-US/firefox/addon/export-cookies/",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 117327,
      "author_name": "bguberfain",
      "author_url": "",
      "post_date": "04/28/2016 15:21:53",
      "content": "<p>It's a little bit tricky, but you can open the Network tab on Chrome developer tools, click to download your train file (on Kaggle website) and watch for the url it uses. Than stop the download, right click the request on the Network tab and choose &quot;Copy as cUrl&quot;. This copies the exact command you can use on a remote machine to download the file (just remember to pipe the output the a file).</p>\n\n<p>I used this method to get files on a cloud computer.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117344,
      "author_name": "chinta",
      "author_url": "",
      "post_date": "04/28/2016 16:17:13",
      "content": "<p>Thanks Bruno. Let me try .</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117352,
      "author_name": "wallych46",
      "author_url": "",
      "post_date": "04/28/2016 17:12:34",
      "content": "<p>Have you seen this?</p>\n\n<p><a href=\"http://www.louisaslett.com/RStudio_AMI/\">http://www.louisaslett.com/RStudio_AMI/</a></p>\n\n<p>I owe Louis money for the time he's saved me.</p>\n\n<p>I start up one of these instances, and while it's installing xgboost etc, I scp the files I need over.  It's quick and easy.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117491,
      "author_name": "chinta",
      "author_url": "",
      "post_date": "04/29/2016 11:25:40",
      "content": "<p>Thanks wally. yes, I did see.</p>\n\n<p>I have uploaded using filezilla. It was easy as well. </p>\n\n<p><a href=\"http://angus.readthedocs.io/en/2014/amazon/transfer-files-between-instance.html\">http://angus.readthedocs.io/en/2014/amazon/transfer-files-between-instance.html</a></p>\n\n<p>when I read test set, it has only 1,875,015 rows, is it not 2,528,243  rows are expected ? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117498,
      "author_name": "wallych46",
      "author_url": "",
      "post_date": "04/29/2016 12:43:38",
      "content": "<p>Yes, 2.5... is what I have.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117613,
      "author_name": "dotank",
      "author_url": "",
      "post_date": "04/29/2016 20:43:19",
      "content": "<p>How is it done if you develop in python? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117659,
      "author_name": "chinta",
      "author_url": "",
      "post_date": "04/30/2016 02:04:37",
      "content": "<p>read the document @ <a href=\"http://angus.readthedocs.io/en/2014/amazon/transfer-files-between-instance.html\">http://angus.readthedocs.io/en/2014/amazon/transfer-files-between-instance.html</a>, it talks about transferring files from local PC to AWS Instance, it has nothing to do with R or python.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117734,
      "author_name": "mafux777",
      "author_url": "",
      "post_date": "04/30/2016 16:05:22",
      "content": "<p>@Bruno - excellent recommendation from you. It took less than one minute to download the .5GB file to my instance on AWS (US-EAST). Also:\nexpedia_train &lt;- fread(&quot;gunzip -c train.csv.gz&quot;, header=TRUE)\nRead 37670293 rows and 24 (of 24) columns from 3.791 GB file in 00:02:02</p>\n\n<p>[quote=Bruno G. do Amaral;117327]</p>\n\n<p>It's a little bit tricky, but you can open the Network tab on Chrome developer tools, click to download your train file (on Kaggle website) and watch for the url it uses. Than stop the download, right click the request on the Network tab and choose &quot;Copy as cUrl&quot;. This copies the exact command you can use on a remote machine to download the file (just remember to pipe the output the a file).</p>\n\n<p>I used this method to get files on a cloud computer.</p>\n\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118302,
      "author_name": "bahramimaryam",
      "author_url": "",
      "post_date": "05/03/2016 07:53:18",
      "content": "<p>Dear Friends\nthis form of reading data does not work on windows system. does it?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118307,
      "author_name": "chinta",
      "author_url": "",
      "post_date": "05/03/2016 08:15:12",
      "content": "<p>@Maryam, I use windows, and I have uploaded using Filezilla, I understand you are asking about the other method, but I did not try that.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119160,
      "author_name": "skylord",
      "author_url": "",
      "post_date": "05/07/2016 17:49:19",
      "content": "<p>If I understood the question correctly, you have downloaded the data on your local machine. And now you were trying to upload it to an AWS instance. Instead why not download it from kaggle directly to the aws instance: </p>\n\n<pre><code>wget --load-cookies=cookies.txt &lt;url to file to download&gt;\n</code></pre>\n\n<p>cookies.txt &gt;&gt; has the cookie data from the kaggle website</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119203,
      "author_name": "jpmiller",
      "author_url": "",
      "post_date": "05/08/2016 01:42:41",
      "content": "<p>Thanks, Skylord -  Each file download took less than 1 second!</p>\n\n<p>Edit: maybe because something isn't quite right...files won't unzip :(</p>\n\n<p>Edit2: It was a problem with my cookies file. Fixed now, and train file took a mere 85 seconds.</p>\n\n<p><img src=\"https://dl.dropboxusercontent.com/u/15775348/ffd.PNG\" alt=\"enter image description here\" title></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119315,
      "author_name": "skylord",
      "author_url": "",
      "post_date": "05/09/2016 06:24:03",
      "content": "<p>JohnM - This was shared by someone in an earlier competitions. Just repeating it over here. Glad you found it useful. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119317,
      "author_name": "alirezaa",
      "author_url": "",
      "post_date": "05/09/2016 06:54:05",
      "content": "<p>You can simply use a cookie-export extension in Chrome. I use the following:\n<a href=\"https://chrome.google.com/webstore/detail/cookietxt-export/lopabhfecdfhgogdbojmaicoicjekelh\">https://chrome.google.com/webstore/detail/cookietxt-export/lopabhfecdfhgogdbojmaicoicjekelh</a></p>\n\n<p>Then just load the cookies as SkyLord commented above.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119357,
      "author_name": "skylord",
      "author_url": "",
      "post_date": "05/09/2016 14:03:12",
      "content": "<p>for a  firefox add-in - \n<a href=\"https://addons.mozilla.org/en-US/firefox/addon/export-cookies/\">https://addons.mozilla.org/en-US/firefox/addon/export-cookies/</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "117318": "Hello,\r\n\r\nI have set up R-Studio on AWS and have been able to upload destinations.csv file, using the option to upload file in R studio, it uploaded fairly quickly. However train.csv upload is taking forever.\r\n\r\nCan you please let me know, if there is a quicker way to upload files.\r\n\r\nThanks,",
    "117327": "It's a little bit tricky, but you can open the Network tab on Chrome developer tools, click to download your train file (on Kaggle website) and watch for the url it uses. Than stop the download, right click the request on the Network tab and choose \"Copy as cUrl\". This copies the exact command you can use on a remote machine to download the file (just remember to pipe the output the a file).\r\n\r\nI used this method to get files on a cloud computer.",
    "117344": "Thanks Bruno. Let me try .",
    "117352": "Have you seen this?\r\n\r\nhttp://www.louisaslett.com/RStudio_AMI/\r\n\r\nI owe Louis money for the time he's saved me.\r\n\r\nI start up one of these instances, and while it's installing xgboost etc, I scp the files I need over.  It's quick and easy.",
    "117491": "Thanks wally. yes, I did see.\r\n\r\nI have uploaded using filezilla. It was easy as well. \r\n\r\nhttp://angus.readthedocs.io/en/2014/amazon/transfer-files-between-instance.html\r\n\r\nwhen I read test set, it has only 1,875,015 rows, is it not 2,528,243  rows are expected ?",
    "117498": "Yes, 2.5... is what I have.",
    "117613": "How is it done if you develop in python?",
    "117659": "read the document @ http://angus.readthedocs.io/en/2014/amazon/transfer-files-between-instance.html, it talks about transferring files from local PC to AWS Instance, it has nothing to do with R or python.",
    "117734": "Bruno - excellent recommendation from you. It took less than one minute to download the .5GB file to my instance on AWS (US-EAST). Also:\r\nexpedia_train <- fread(\"gunzip -c train.csv.gz\", header=TRUE)\r\nRead 37670293 rows and 24 (of 24) columns from 3.791 GB file in 00:02:02\r\n\r\n[quote=Bruno G. do Amaral;117327]\r\n\r\nIt's a little bit tricky, but you can open the Network tab on Chrome developer tools, click to download your train file (on Kaggle website) and watch for the url it uses. Than stop the download, right click the request on the Network tab and choose \"Copy as cUrl\". This copies the exact command you can use on a remote machine to download the file (just remember to pipe the output the a file).\r\n\r\nI used this method to get files on a cloud computer.\r\n\r\n[/quote]",
    "118302": "Dear Friends\r\nthis form of reading data does not work on windows system. does it?",
    "118307": "Maryam, I use windows, and I have uploaded using Filezilla, I understand you are asking about the other method, but I did not try that.",
    "119160": "If I understood the question correctly, you have downloaded the data on your local machine. And now you were trying to upload it to an AWS instance. Instead why not download it from kaggle directly to the aws instance: \r\n\r\n    wget --load-cookies=cookies.txt <url to file to download>\r\n\r\ncookies.txt >> has the cookie data from the kaggle website",
    "119203": "Thanks, Skylord -  Each file download took less than 1 second!\r\n\r\nEdit: maybe because something isn't quite right...files won't unzip :(\r\n\r\nEdit2: It was a problem with my cookies file. Fixed now, and train file took a mere 85 seconds.\r\n\r\n![enter image description here][1]\r\n\r\n\r\n  [1]: https://dl.dropboxusercontent.com/u/15775348/ffd.PNG",
    "119315": "JohnM - This was shared by someone in an earlier competitions. Just repeating it over here. Glad you found it useful.",
    "119317": "You can simply use a cookie-export extension in Chrome. I use the following:\r\nhttps://chrome.google.com/webstore/detail/cookietxt-export/lopabhfecdfhgogdbojmaicoicjekelh\r\n\r\nThen just load the cookies as SkyLord commented above.",
    "119357": "for a  firefox add-in - \r\nhttps://addons.mozilla.org/en-US/firefox/addon/export-cookies/"
  },
  "source": "meta"
}