{
  "id": 155624,
  "title": "Download individual folders via python",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/155624",
  "author_name": "",
  "post_date": "2020-06-02T11:34:22.795297900Z",
  "votes": 10,
  "comment_count": 3,
  "views": 0,
  "content": "<p>### Hi all. A lot of times, we need to download individual folders from a competition because we do not need all the files. </p>\n\n<h3>The API as of now does not support this so the only way I used to do this was to download each folder to my local system and then upload the same to the cloud server. Long process ☹️</h3>\n\n<h3>Recently while experimenting, I stumbled upon a neat little trick to get individual folders into my cloud server directly. Listing the steps below so that it may help anyone who needs it 😄</h3>\n\n<p><strong>Step 1:</strong>\nIn this competition, I wanted only the JPEG folder. So click the JPEG folder and on the right side you will see a download link which looks like this: \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4807296%2F0af736dfed3ffdd25d4192cf8c009e38%2F1.JPG?generation=1591096910072006&amp;alt=media\" alt=\"\"></p>\n\n<p>Click on the link to start the download. After the download starts, pause it. Click on the arrow (circled) and click pause. </p>\n\n<p><strong>Step 2</strong>\nOpen up downloads. On chrome, the shortcut is Ctrl + J. There, copy the download link of the file being downloaded. Image below for reference \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4807296%2Fddd65cc3738b64d767bae2e151fc540f%2F2.JPG?generation=1591097004191209&amp;alt=media\" alt=\"\"></p>\n\n<p><strong>Step 3</strong>\nWe will now download the file using <code>wget</code></p>\n\n<p>Open up your editor, and use the following code to download the file </p>\n\n<p><code>import wget</code>\n <code>url = {*Paste copied url here*}</code>\n<code>wget.download(url, 'data_path/train.zip')</code></p>\n\n<p><em>A key point to note here is that you have to specify the download path as <code>path/filename.zip</code>. Just providing the path won't work</em>. You can specify any name here. Doesn't have to be the same as Kaggle's name (jpeg.zip). Image below for reference \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4807296%2Ff453f6b91dff24ba119ab627da60abfd%2F3.JPG?generation=1591097333265574&amp;alt=media\" alt=\"\"></p>\n\n<p>That's all. Now you can extract the file and proceed to train. Same steps for any other folder. Hopefully this helps some fellow practitioners out there 😄 </p>",
  "messages": [
    {
      "id": "871465",
      "postDate": "06/02/2020 11:34:22",
      "content": "<p>### Hi all. A lot of times, we need to download individual folders from a competition because we do not need all the files. </p>\n\n<h3>The API as of now does not support this so the only way I used to do this was to download each folder to my local system and then upload the same to the cloud server. Long process ☹️</h3>\n\n<h3>Recently while experimenting, I stumbled upon a neat little trick to get individual folders into my cloud server directly. Listing the steps below so that it may help anyone who needs it 😄</h3>\n\n<p><strong>Step 1:</strong>\nIn this competition, I wanted only the JPEG folder. So click the JPEG folder and on the right side you will see a download link which looks like this: \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4807296%2F0af736dfed3ffdd25d4192cf8c009e38%2F1.JPG?generation=1591096910072006&amp;alt=media\" alt=\"\"></p>\n\n<p>Click on the link to start the download. After the download starts, pause it. Click on the arrow (circled) and click pause. </p>\n\n<p><strong>Step 2</strong>\nOpen up downloads. On chrome, the shortcut is Ctrl + J. There, copy the download link of the file being downloaded. Image below for reference \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4807296%2Fddd65cc3738b64d767bae2e151fc540f%2F2.JPG?generation=1591097004191209&amp;alt=media\" alt=\"\"></p>\n\n<p><strong>Step 3</strong>\nWe will now download the file using <code>wget</code></p>\n\n<p>Open up your editor, and use the following code to download the file </p>\n\n<p><code>import wget</code>\n <code>url = {*Paste copied url here*}</code>\n<code>wget.download(url, 'data_path/train.zip')</code></p>\n\n<p><em>A key point to note here is that you have to specify the download path as <code>path/filename.zip</code>. Just providing the path won't work</em>. You can specify any name here. Doesn't have to be the same as Kaggle's name (jpeg.zip). Image below for reference \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4807296%2Ff453f6b91dff24ba119ab627da60abfd%2F3.JPG?generation=1591097333265574&amp;alt=media\" alt=\"\"></p>\n\n<p>That's all. Now you can extract the file and proceed to train. Same steps for any other folder. Hopefully this helps some fellow practitioners out there 😄 </p>",
      "rawMarkdown": "### Hi all. A lot of times, we need to download individual folders from a competition because we do not need all the files. \n###The API as of now does not support this so the only way I used to do this was to download each folder to my local system and then upload the same to the cloud server. Long process ☹️ \n\n###Recently while experimenting, I stumbled upon a neat little trick to get individual folders into my cloud server directly. Listing the steps below so that it may help anyone who needs it 😄 \n\n**Step 1:**\nIn this competition, I wanted only the JPEG folder. So click the JPEG folder and on the right side you will see a download link which looks like this: \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4807296%2F0af736dfed3ffdd25d4192cf8c009e38%2F1.JPG?generation=1591096910072006&amp;alt=media)\n\nClick on the link to start the download. After the download starts, pause it. Click on the arrow (circled) and click pause. \n\n**Step 2**\nOpen up downloads. On chrome, the shortcut is Ctrl + J. There, copy the download link of the file being downloaded. Image below for reference \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4807296%2Fddd65cc3738b64d767bae2e151fc540f%2F2.JPG?generation=1591097004191209&amp;alt=media)\n\n**Step 3**\nWe will now download the file using `wget`\n\nOpen up your editor, and use the following code to download the file \n\n`import wget`\n `url = {*Paste copied url here*}`\n` wget.download(url, 'data_path/train.zip')`\n\n*A key point to note here is that you have to specify the download path as `path/filename.zip`. Just providing the path won't work*. You can specify any name here. Doesn't have to be the same as Kaggle's name (jpeg.zip). Image below for reference \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4807296%2Ff453f6b91dff24ba119ab627da60abfd%2F3.JPG?generation=1591097333265574&amp;alt=media)\n\nThat's all. Now you can extract the file and proceed to train. Same steps for any other folder. Hopefully this helps some fellow practitioners out there 😄",
      "votes": null
    },
    {
      "id": "871679",
      "postDate": "06/02/2020 15:09:11",
      "content": "<p>It works, thank you!</p>",
      "rawMarkdown": "It works, thank you!",
      "votes": null
    },
    {
      "id": "871730",
      "postDate": "06/02/2020 15:45:49",
      "content": "<p>I wanted to add something to this, hopefully it would help beginners.</p>\n\n<p>You can <strong>resume</strong> broken Chrome downloads from Firefox. Here are some helpful links: </p>\n\n<ul>\n<li><a href=\"https://www.ubergizmo.com/how-to/how-to-resume-chrome-download-using-firefox/\">UberGizmo Guide</a></li>\n<li><a href=\"https://support.mozilla.org/en-US/questions/1220917\">Mozilla Forums</a></li>\n<li><a href=\"https://lifehacker.com/resume-a-failed-chrome-download-with-firefox-1655246429\">Lifehacker Article</a></li>\n</ul>",
      "rawMarkdown": "I wanted to add something to this, hopefully it would help beginners.\n\nYou can **resume** broken Chrome downloads from Firefox. Here are some helpful links: \n\n- [UberGizmo Guide](https://www.ubergizmo.com/how-to/how-to-resume-chrome-download-using-firefox/)\n- [Mozilla Forums](https://support.mozilla.org/en-US/questions/1220917)\n- [Lifehacker Article](https://lifehacker.com/resume-a-failed-chrome-download-with-firefox-1655246429)",
      "votes": null
    },
    {
      "id": "875073",
      "postDate": "06/05/2020 14:03:23",
      "content": "<p>Alternatively, instead of using Python you could use wget in LInux terminal itself:\n<code>wget -O \"./zip_name.zip\" \"url_to_download\"</code></p>\n\n<p>Make sure to add the double quotes to include long urls.</p>",
      "rawMarkdown": "Alternatively, instead of using Python you could use wget in LInux terminal itself:\n`wget -O \"./zip_name.zip\" \"url_to_download\"`\n\nMake sure to add the double quotes to include long urls.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 871679,
      "author_name": "maximofn",
      "author_url": "",
      "post_date": "06/02/2020 15:09:11",
      "content": "<p>It works, thank you!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 871730,
      "author_name": "aadhavvignesh",
      "author_url": "",
      "post_date": "06/02/2020 15:45:49",
      "content": "<p>I wanted to add something to this, hopefully it would help beginners.</p>\n\n<p>You can <strong>resume</strong> broken Chrome downloads from Firefox. Here are some helpful links: </p>\n\n<ul>\n<li><a href=\"https://www.ubergizmo.com/how-to/how-to-resume-chrome-download-using-firefox/\">UberGizmo Guide</a></li>\n<li><a href=\"https://support.mozilla.org/en-US/questions/1220917\">Mozilla Forums</a></li>\n<li><a href=\"https://lifehacker.com/resume-a-failed-chrome-download-with-firefox-1655246429\">Lifehacker Article</a></li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 875073,
      "author_name": "deep1845kaggle",
      "author_url": "",
      "post_date": "06/05/2020 14:03:23",
      "content": "<p>Alternatively, instead of using Python you could use wget in LInux terminal itself:\n<code>wget -O \"./zip_name.zip\" \"url_to_download\"</code></p>\n\n<p>Make sure to add the double quotes to include long urls.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "871465": "### Hi all. A lot of times, we need to download individual folders from a competition because we do not need all the files. \n###The API as of now does not support this so the only way I used to do this was to download each folder to my local system and then upload the same to the cloud server. Long process ☹️ \n\n###Recently while experimenting, I stumbled upon a neat little trick to get individual folders into my cloud server directly. Listing the steps below so that it may help anyone who needs it 😄 \n\n**Step 1:**\nIn this competition, I wanted only the JPEG folder. So click the JPEG folder and on the right side you will see a download link which looks like this: \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4807296%2F0af736dfed3ffdd25d4192cf8c009e38%2F1.JPG?generation=1591096910072006&amp;alt=media)\n\nClick on the link to start the download. After the download starts, pause it. Click on the arrow (circled) and click pause. \n\n**Step 2**\nOpen up downloads. On chrome, the shortcut is Ctrl + J. There, copy the download link of the file being downloaded. Image below for reference \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4807296%2Fddd65cc3738b64d767bae2e151fc540f%2F2.JPG?generation=1591097004191209&amp;alt=media)\n\n**Step 3**\nWe will now download the file using `wget`\n\nOpen up your editor, and use the following code to download the file \n\n`import wget`\n `url = {*Paste copied url here*}`\n` wget.download(url, 'data_path/train.zip')`\n\n*A key point to note here is that you have to specify the download path as `path/filename.zip`. Just providing the path won't work*. You can specify any name here. Doesn't have to be the same as Kaggle's name (jpeg.zip). Image below for reference \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4807296%2Ff453f6b91dff24ba119ab627da60abfd%2F3.JPG?generation=1591097333265574&amp;alt=media)\n\nThat's all. Now you can extract the file and proceed to train. Same steps for any other folder. Hopefully this helps some fellow practitioners out there 😄",
    "871679": "It works, thank you!",
    "871730": "I wanted to add something to this, hopefully it would help beginners.\n\nYou can **resume** broken Chrome downloads from Firefox. Here are some helpful links: \n\n- [UberGizmo Guide](https://www.ubergizmo.com/how-to/how-to-resume-chrome-download-using-firefox/)\n- [Mozilla Forums](https://support.mozilla.org/en-US/questions/1220917)\n- [Lifehacker Article](https://lifehacker.com/resume-a-failed-chrome-download-with-firefox-1655246429)",
    "875073": "Alternatively, instead of using Python you could use wget in LInux terminal itself:\n`wget -O \"./zip_name.zip\" \"url_to_download\"`\n\nMake sure to add the double quotes to include long urls."
  },
  "source": "meta"
}