{
  "id": 154364,
  "title": "Kaggle API jpeg only downloads",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/154364",
  "author_name": "olivier",
  "post_date": "2020-05-28T05:43:26.418000",
  "votes": 22,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Hi KAGGLE team, using kaggle API I have been trying to download only partial data without luck.</p>\n\n<p>I tried this but did not work :( \n<code>\nkaggle c download siim-isic-melanoma-classification -f jpeg.zip\n</code></p>\n\n<p>Is there a specific archive to avoid having to download 100GB of data ?</p>\n\n<p>Or do I need to download the jpeg files one by one using train and test csv files ?</p>\n\n<p>Thanks for your help.</p>",
  "messages": [
    {
      "id": 864654,
      "postDate": "2020-05-28T05:43:26.417Z",
      "content": "<p>Hi KAGGLE team, using kaggle API I have been trying to download only partial data without luck.</p>\n\n<p>I tried this but did not work :( \n<code>\nkaggle c download siim-isic-melanoma-classification -f jpeg.zip\n</code></p>\n\n<p>Is there a specific archive to avoid having to download 100GB of data ?</p>\n\n<p>Or do I need to download the jpeg files one by one using train and test csv files ?</p>\n\n<p>Thanks for your help.</p>",
      "rawMarkdown": "Hi KAGGLE team, using kaggle API I have been trying to download only partial data without luck.\n\nI tried this but did not work :( \n```\nkaggle c download siim-isic-melanoma-classification -f jpeg.zip\n```\n\nIs there a specific archive to avoid having to download 100GB of data ?\n\nOr do I need to download the jpeg files one by one using train and test csv files ?\n\nThanks for your help.",
      "votes": 20
    },
    {
      "id": 865584,
      "postDate": "2020-05-28T18:07:59.947Z",
      "content": "<p>Yes, unfortunately, it's not possible to download a sub-folder via API at this time. They'd have to be downloaded one-by-one. As Pavel points out, it is possible to download a sub-folder via the data explorer. Sorry, we recognize this is not ideal and hope to resolve it in the future.</p>",
      "rawMarkdown": "Yes, unfortunately, it's not possible to download a sub-folder via API at this time. They'd have to be downloaded one-by-one. As Pavel points out, it is possible to download a sub-folder via the data explorer. Sorry, we recognize this is not ideal and hope to resolve it in the future.",
      "votes": 7,
      "replies": [
        {
          "id": 868322,
          "postDate": "2020-05-31T06:30:36.790Z",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> Could you please upgrade the kaggle api so that we can download only the particular files that we need to use. Because some of participants like me doesn't have the required storage space to download that much data to participate in this competition.Hope kaggle dev team upgrades this api quickly so that we can also join in this one </p>",
          "rawMarkdown": "@juliaelliott Could you please upgrade the kaggle api so that we can download only the particular files that we need to use. Because some of participants like me doesn't have the required storage space to download that much data to participate in this competition.Hope kaggle dev team upgrades this api quickly so that we can also join in this one ",
          "votes": 4
        }
      ]
    },
    {
      "id": 870179,
      "postDate": "2020-06-01T14:42:34.813Z",
      "content": "<p><a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121194\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121194</a></p>\n\n<p>this might help you</p>",
      "rawMarkdown": "https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121194\n\nthis might help you",
      "votes": 1
    },
    {
      "id": 866926,
      "postDate": "2020-05-29T19:42:45.907Z",
      "content": "<p>You can get the list of files through the api and then download specific files. Here's an example for getting just the csv and tfrecord files.</p>\n\n<p>``` python\nfile_list = kaggle.api.competition_list_files('siim-isic-melanoma-classification')\ncsv_files = [str(f) for f in file_list if 'csv' in str(f)]\ntfrec_files = [str(f) for f in file_list if 'tfrec' in str(f)]</p>\n\n<p>for filename in tqdm(csv_files + tfrec_files):\n    kaggle.api.competition_download_file('siim-isic-melanoma-classification',\n                                         filename,\n                                         './data/melanoma',\n                                         quiet=True)\n```</p>",
      "rawMarkdown": "You can get the list of files through the api and then download specific files. Here's an example for getting just the csv and tfrecord files.\n\n``` python\nfile_list = kaggle.api.competition_list_files('siim-isic-melanoma-classification')\ncsv_files = [str(f) for f in file_list if 'csv' in str(f)]\ntfrec_files = [str(f) for f in file_list if 'tfrec' in str(f)]\n\nfor filename in tqdm(csv_files + tfrec_files):\n    kaggle.api.competition_download_file('siim-isic-melanoma-classification',\n                                         filename,\n                                         './data/melanoma',\n                                         quiet=True)\n```",
      "votes": 1,
      "replies": [
        {
          "id": 867122,
          "postDate": "2020-05-30T02:44:43.613Z",
          "content": "<p>This actually ended up not working. For some reason the api couldn't find all of the individual files. I was able to download the 15 test tfrecord files, but only the first four train files.</p>",
          "rawMarkdown": "This actually ended up not working. For some reason the api couldn't find all of the individual files. I was able to download the 15 test tfrecord files, but only the first four train files.",
          "votes": 4
        }
      ]
    },
    {
      "id": 865001,
      "postDate": "2020-05-28T10:23:27.247Z",
      "content": "<p>I was able to download it manually jpeg.zip. If you select the \"jpeg\" folder (in \"Data Explorer\") on the Data tab and click the download button. Downloaded with a very long delay after several gigabytes ... but completed successfully (~31Gb)</p>",
      "rawMarkdown": "I was able to download it manually jpeg.zip. If you select the \"jpeg\" folder (in \"Data Explorer\") on the Data tab and click the download button. Downloaded with a very long delay after several gigabytes ... but completed successfully (~31Gb)",
      "votes": 2,
      "replies": [
        {
          "id": 865022,
          "postDate": "2020-05-28T10:35:50.467Z",
          "content": "<p>Thanks <a href=\"/sapr3s\">@sapr3s</a> from a GUI yes you can but I'm trying to do that on an ubuntu server from the CLI :) </p>",
          "rawMarkdown": "Thanks @sapr3s from a GUI yes you can but I'm trying to do that on an ubuntu server from the CLI :) "
        }
      ]
    },
    {
      "id": 890426,
      "postDate": "2020-06-17T14:02:24.750Z",
      "content": "<p><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/159456\">This</a> might help you. I have successfully downloaded jpeg folder on colab by following this.</p>",
      "rawMarkdown": "[This](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/159456) might help you. I have successfully downloaded jpeg folder on colab by following this."
    },
    {
      "id": 864972,
      "postDate": "2020-05-28T10:06:41.923Z",
      "content": "<p>Well the csv file contains the unique ids, so just need to get this, add <code>.jpg</code> and issue the right <code>kaggle c download -f file_name</code> command. That's only around 45K files to download this way...</p>",
      "rawMarkdown": "Well the csv file contains the unique ids, so just need to get this, add `.jpg` and issue the right `kaggle c download -f file_name` command. That's only around 45K files to download this way..."
    },
    {
      "id": 864965,
      "postDate": "2020-05-28T10:00:43.477Z",
      "content": "<p>how are you going to download the jpeg files from a csv file??</p>",
      "rawMarkdown": "how are you going to download the jpeg files from a csv file??\n",
      "replies": [
        {
          "id": 867107,
          "postDate": "2020-05-30T02:26:30.283Z",
          "content": "<p>Not actually downloading the jpg from csv but using csv to extract the unique ids/filename of the jpg images</p>",
          "rawMarkdown": "Not actually downloading the jpg from csv but using csv to extract the unique ids/filename of the jpg images",
          "votes": 1
        }
      ]
    },
    {
      "id": 866475,
      "postDate": "2020-05-29T12:46:31.417Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 865584,
      "author_name": "Julia Elliott",
      "author_url": "",
      "post_date": "2020-05-28T18:07:59.947000",
      "content": "<p>Yes, unfortunately, it's not possible to download a sub-folder via API at this time. They'd have to be downloaded one-by-one. As Pavel points out, it is possible to download a sub-folder via the data explorer. Sorry, we recognize this is not ideal and hope to resolve it in the future.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 868322,
          "author_name": "MhdSharuk",
          "author_url": "",
          "post_date": "2020-05-31T06:30:36.790000",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> Could you please upgrade the kaggle api so that we can download only the particular files that we need to use. Because some of participants like me doesn't have the required storage space to download that much data to participate in this competition.Hope kaggle dev team upgrades this api quickly so that we can also join in this one </p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 870179,
      "author_name": "rajitha",
      "author_url": "",
      "post_date": "2020-06-01T14:42:34.813000",
      "content": "<p><a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121194\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121194</a></p>\n\n<p>this might help you</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 866926,
      "author_name": "Caleb",
      "author_url": "",
      "post_date": "2020-05-29T19:42:45.907000",
      "content": "<p>You can get the list of files through the api and then download specific files. Here's an example for getting just the csv and tfrecord files.</p>\n\n<p>``` python\nfile_list = kaggle.api.competition_list_files('siim-isic-melanoma-classification')\ncsv_files = [str(f) for f in file_list if 'csv' in str(f)]\ntfrec_files = [str(f) for f in file_list if 'tfrec' in str(f)]</p>\n\n<p>for filename in tqdm(csv_files + tfrec_files):\n    kaggle.api.competition_download_file('siim-isic-melanoma-classification',\n                                         filename,\n                                         './data/melanoma',\n                                         quiet=True)\n```</p>",
      "votes": 1,
      "replies": [
        {
          "id": 867122,
          "author_name": "Caleb",
          "author_url": "",
          "post_date": "2020-05-30T02:44:43.613000",
          "content": "<p>This actually ended up not working. For some reason the api couldn't find all of the individual files. I was able to download the 15 test tfrecord files, but only the first four train files.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 865001,
      "author_name": "Pavel Orlov",
      "author_url": "",
      "post_date": "2020-05-28T10:23:27.247000",
      "content": "<p>I was able to download it manually jpeg.zip. If you select the \"jpeg\" folder (in \"Data Explorer\") on the Data tab and click the download button. Downloaded with a very long delay after several gigabytes ... but completed successfully (~31Gb)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 865022,
          "author_name": "olivier",
          "author_url": "",
          "post_date": "2020-05-28T10:35:50.467000",
          "content": "<p>Thanks <a href=\"/sapr3s\">@sapr3s</a> from a GUI yes you can but I'm trying to do that on an ubuntu server from the CLI :) </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 890426,
      "author_name": "Kishan Joshi",
      "author_url": "",
      "post_date": "2020-06-17T14:02:24.750000",
      "content": "<p><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/159456\">This</a> might help you. I have successfully downloaded jpeg folder on colab by following this.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 864972,
      "author_name": "olivier",
      "author_url": "",
      "post_date": "2020-05-28T10:06:41.923000",
      "content": "<p>Well the csv file contains the unique ids, so just need to get this, add <code>.jpg</code> and issue the right <code>kaggle c download -f file_name</code> command. That's only around 45K files to download this way...</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 864965,
      "author_name": "MhdSharuk",
      "author_url": "",
      "post_date": "2020-05-28T10:00:43.477000",
      "content": "<p>how are you going to download the jpeg files from a csv file??</p>",
      "votes": 0,
      "replies": [
        {
          "id": 867107,
          "author_name": "Anonymous",
          "author_url": "",
          "post_date": "2020-05-30T02:26:30.283000",
          "content": "<p>Not actually downloading the jpg from csv but using csv to extract the unique ids/filename of the jpg images</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 866475,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-05-29T12:46:31.417000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "864654": "Hi KAGGLE team, using kaggle API I have been trying to download only partial data without luck.\n\nI tried this but did not work :( \n```\nkaggle c download siim-isic-melanoma-classification -f jpeg.zip\n```\n\nIs there a specific archive to avoid having to download 100GB of data ?\n\nOr do I need to download the jpeg files one by one using train and test csv files ?\n\nThanks for your help.",
    "865584": "Yes, unfortunately, it's not possible to download a sub-folder via API at this time. They'd have to be downloaded one-by-one. As Pavel points out, it is possible to download a sub-folder via the data explorer. Sorry, we recognize this is not ideal and hope to resolve it in the future.",
    "870179": "https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121194\n\nthis might help you",
    "866926": "You can get the list of files through the api and then download specific files. Here's an example for getting just the csv and tfrecord files.\n\n``` python\nfile_list = kaggle.api.competition_list_files('siim-isic-melanoma-classification')\ncsv_files = [str(f) for f in file_list if 'csv' in str(f)]\ntfrec_files = [str(f) for f in file_list if 'tfrec' in str(f)]\n\nfor filename in tqdm(csv_files + tfrec_files):\n    kaggle.api.competition_download_file('siim-isic-melanoma-classification',\n                                         filename,\n                                         './data/melanoma',\n                                         quiet=True)\n```",
    "865001": "I was able to download it manually jpeg.zip. If you select the \"jpeg\" folder (in \"Data Explorer\") on the Data tab and click the download button. Downloaded with a very long delay after several gigabytes ... but completed successfully (~31Gb)",
    "890426": "[This](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/159456) might help you. I have successfully downloaded jpeg folder on colab by following this.",
    "864972": "Well the csv file contains the unique ids, so just need to get this, add `.jpg` and issue the right `kaggle c download -f file_name` command. That's only around 45K files to download this way...",
    "864965": "how are you going to download the jpeg files from a csv file??\n",
    "866475": ""
  }
}