{
  "id": 16409,
  "title": "Downloading Data",
  "url": "/competitions/dato-native/discussion/16409",
  "author_name": "",
  "post_date": "2015-09-10T13:20:17.500Z",
  "votes": null,
  "comment_count": 6,
  "views": 784,
  "content": "<p>Hi I am trying to download the data on a EC2 machine . I tried the link : <a href=\"https://www.kaggle.com/forums/f/15/kaggle-forum/t/6604/downloading-data-via-command-line\">https://www.kaggle.com/forums/f/15/kaggle-forum/t/6604/downloading-data-via-command-line</a> but it is not helping . Anyone who has a clue ???</p>",
  "messages": [
    {
      "id": "92006",
      "postDate": "09/10/2015 13:20:17",
      "content": "<p>Hi I am trying to download the data on a EC2 machine . I tried the link : <a href=\"https://www.kaggle.com/forums/f/15/kaggle-forum/t/6604/downloading-data-via-command-line\">https://www.kaggle.com/forums/f/15/kaggle-forum/t/6604/downloading-data-via-command-line</a> but it is not helping . Anyone who has a clue ???</p>",
      "rawMarkdown": "Hi I am trying to download the data on a EC2 machine . I tried the link : https://www.kaggle.com/forums/f/15/kaggle-forum/t/6604/downloading-data-via-command-line but it is not helping . Anyone who has a clue ???",
      "votes": null
    },
    {
      "id": "92009",
      "postDate": "09/10/2015 15:31:52",
      "content": "<p>I download it to my local machine into a Dropbox folder, and then have Dropbox on the AWS machine.</p>\n\n<p>The main issue there is that as scripts output large files, Dropbox will want to start updating it back to itself, and it can compete for resources on the machine. You can turn it off, but then need to remember to turn it back on again. \nHaving this also gives you some versioning and rollback ability too in cases of fat fingering something and losing a file.</p>",
      "rawMarkdown": "I download it to my local machine into a Dropbox folder, and then have Dropbox on the AWS machine.\r\n\r\nThe main issue there is that as scripts output large files, Dropbox will want to start updating it back to itself, and it can compete for resources on the machine. You can turn it off, but then need to remember to turn it back on again. \r\nHaving this also gives you some versioning and rollback ability too in cases of fat fingering something and losing a file.",
      "votes": null
    },
    {
      "id": "92019",
      "postDate": "09/10/2015 17:22:02",
      "content": "<p>You can start downloading locally with chrome, pause, copy the link it downloads from and use <code>curl -O</code> on EC2. </p>",
      "rawMarkdown": "You can start downloading locally with chrome, pause, copy the link it downloads from and use `curl -O` on EC2.",
      "votes": null
    },
    {
      "id": "92036",
      "postDate": "09/10/2015 19:48:53",
      "content": "<p>@Dmitry It gives  resource not found error</p>",
      "rawMarkdown": "Dmitry It gives  resource not found error",
      "votes": null
    },
    {
      "id": "92037",
      "postDate": "09/10/2015 20:10:19",
      "content": "<p>And you have definitely clicked on it in a browser to activate the javascript thing (the default link) that makes a div popup asking you to confirm if you are a team or not, and then you are using the link after that?\nOr are you using the link the first time visiting the page?</p>",
      "rawMarkdown": "And you have definitely clicked on it in a browser to activate the javascript thing (the default link) that makes a div popup asking you to confirm if you are a team or not, and then you are using the link after that?\r\nOr are you using the link the first time visiting the page?",
      "votes": null
    },
    {
      "id": "92038",
      "postDate": "09/10/2015 20:12:55",
      "content": "<p>no I have looked at the page multiple times before downloading data. Also , Though I am a part of team but still it never asked me .</p>",
      "rawMarkdown": "no I have looked at the page multiple times before downloading data. Also , Though I am a part of team but still it never asked me .",
      "votes": null
    },
    {
      "id": "92247",
      "postDate": "09/12/2015 13:43:57",
      "content": "<p>I just did this today:</p>\n\n<ul>\n<li><strong>Use chrome</strong>, inspect the page, go to network tab.</li>\n<li>Click download, say, 1.zip</li>\n<li>You will see a request that downloads 1.zip, ignore this one.</li>\n<li>It will spawn another request that downloads the real 1.zip. </li>\n<li>Right click the request, then copy the curl command.</li>\n<li>Paste it to EC2 terminal and run. Remember to redirect the output to a file (e.g. curl xxxxx &gt; 1.zip)</li>\n</ul>\n\n<p>Let me know if it works for you.</p>",
      "rawMarkdown": "I just did this today:\r\n\r\n- **Use chrome**, inspect the page, go to network tab.\r\n- Click download, say, 1.zip\r\n- You will see a request that downloads 1.zip, ignore this one.\r\n- It will spawn another request that downloads the real 1.zip. \r\n- Right click the request, then copy the curl command.\r\n- Paste it to EC2 terminal and run. Remember to redirect the output to a file (e.g. curl xxxxx > 1.zip)\r\n\r\nLet me know if it works for you.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 92009,
      "author_name": "omgponies",
      "author_url": "",
      "post_date": "09/10/2015 15:31:52",
      "content": "<p>I download it to my local machine into a Dropbox folder, and then have Dropbox on the AWS machine.</p>\n\n<p>The main issue there is that as scripts output large files, Dropbox will want to start updating it back to itself, and it can compete for resources on the machine. You can turn it off, but then need to remember to turn it back on again. \nHaving this also gives you some versioning and rollback ability too in cases of fat fingering something and losing a file.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92019,
      "author_name": "dulyanov",
      "author_url": "",
      "post_date": "09/10/2015 17:22:02",
      "content": "<p>You can start downloading locally with chrome, pause, copy the link it downloads from and use <code>curl -O</code> on EC2. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92036,
      "author_name": "anirudhkala",
      "author_url": "",
      "post_date": "09/10/2015 19:48:53",
      "content": "<p>@Dmitry It gives  resource not found error</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92037,
      "author_name": "omgponies",
      "author_url": "",
      "post_date": "09/10/2015 20:10:19",
      "content": "<p>And you have definitely clicked on it in a browser to activate the javascript thing (the default link) that makes a div popup asking you to confirm if you are a team or not, and then you are using the link after that?\nOr are you using the link the first time visiting the page?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92038,
      "author_name": "anirudhkala",
      "author_url": "",
      "post_date": "09/10/2015 20:12:55",
      "content": "<p>no I have looked at the page multiple times before downloading data. Also , Though I am a part of team but still it never asked me .</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 92247,
      "author_name": "goodwind",
      "author_url": "",
      "post_date": "09/12/2015 13:43:57",
      "content": "<p>I just did this today:</p>\n\n<ul>\n<li><strong>Use chrome</strong>, inspect the page, go to network tab.</li>\n<li>Click download, say, 1.zip</li>\n<li>You will see a request that downloads 1.zip, ignore this one.</li>\n<li>It will spawn another request that downloads the real 1.zip. </li>\n<li>Right click the request, then copy the curl command.</li>\n<li>Paste it to EC2 terminal and run. Remember to redirect the output to a file (e.g. curl xxxxx &gt; 1.zip)</li>\n</ul>\n\n<p>Let me know if it works for you.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "92006": "Hi I am trying to download the data on a EC2 machine . I tried the link : https://www.kaggle.com/forums/f/15/kaggle-forum/t/6604/downloading-data-via-command-line but it is not helping . Anyone who has a clue ???",
    "92009": "I download it to my local machine into a Dropbox folder, and then have Dropbox on the AWS machine.\r\n\r\nThe main issue there is that as scripts output large files, Dropbox will want to start updating it back to itself, and it can compete for resources on the machine. You can turn it off, but then need to remember to turn it back on again. \r\nHaving this also gives you some versioning and rollback ability too in cases of fat fingering something and losing a file.",
    "92019": "You can start downloading locally with chrome, pause, copy the link it downloads from and use `curl -O` on EC2.",
    "92036": "Dmitry It gives  resource not found error",
    "92037": "And you have definitely clicked on it in a browser to activate the javascript thing (the default link) that makes a div popup asking you to confirm if you are a team or not, and then you are using the link after that?\r\nOr are you using the link the first time visiting the page?",
    "92038": "no I have looked at the page multiple times before downloading data. Also , Though I am a part of team but still it never asked me .",
    "92247": "I just did this today:\r\n\r\n- **Use chrome**, inspect the page, go to network tab.\r\n- Click download, say, 1.zip\r\n- You will see a request that downloads 1.zip, ignore this one.\r\n- It will spawn another request that downloads the real 1.zip. \r\n- Right click the request, then copy the curl command.\r\n- Paste it to EC2 terminal and run. Remember to redirect the output to a file (e.g. curl xxxxx > 1.zip)\r\n\r\nLet me know if it works for you."
  },
  "source": "meta"
}