{
  "id": 523419,
  "title": "Efficient Methods for Downloading Large Datasets (15GB+) in Smaller Chunks?",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/523419",
  "author_name": "",
  "post_date": "2024-08-01T01:41:24.855022200Z",
  "votes": 4,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi Folks!</p>\n<p>I hope you're all doing well. I'm currently working on this project but the dataset exceeds 34GB in size. Given the substantial size, downloading the entire dataset in one go is not feasible due to time constraints and bandwidth limitations.</p>\n<p>I'd like to ask for advice on a couple of points:</p>\n<p>Is there an alternative way to download the dataset in smaller chunks (e.g., in MB)? This would help manage the download process more efficiently and avoid potential interruptions or failures during the download.</p>\n<p>Are there any tools or scripts that you recommend for this purpose? Perhaps something that can automate the process and resume downloading in case of any disruptions?</p>\n<p>Any insights, suggestions, or examples would be greatly appreciated. If you've faced a similar situation and found a workaround, please share your experience!</p>\n<p>Thank you in advance for your help!</p>",
  "messages": [
    {
      "id": "2942650",
      "postDate": "08/01/2024 01:41:24",
      "content": "<p>Hi Folks!</p>\n<p>I hope you're all doing well. I'm currently working on this project but the dataset exceeds 34GB in size. Given the substantial size, downloading the entire dataset in one go is not feasible due to time constraints and bandwidth limitations.</p>\n<p>I'd like to ask for advice on a couple of points:</p>\n<p>Is there an alternative way to download the dataset in smaller chunks (e.g., in MB)? This would help manage the download process more efficiently and avoid potential interruptions or failures during the download.</p>\n<p>Are there any tools or scripts that you recommend for this purpose? Perhaps something that can automate the process and resume downloading in case of any disruptions?</p>\n<p>Any insights, suggestions, or examples would be greatly appreciated. If you've faced a similar situation and found a workaround, please share your experience!</p>\n<p>Thank you in advance for your help!</p>",
      "rawMarkdown": "Hi Folks!\n\nI hope you're all doing well. I'm currently working on this project but the dataset exceeds 34GB in size. Given the substantial size, downloading the entire dataset in one go is not feasible due to time constraints and bandwidth limitations.\n\nI'd like to ask for advice on a couple of points:\n\nIs there an alternative way to download the dataset in smaller chunks (e.g., in MB)? This would help manage the download process more efficiently and avoid potential interruptions or failures during the download.\n\nAre there any tools or scripts that you recommend for this purpose? Perhaps something that can automate the process and resume downloading in case of any disruptions?\n\nAny insights, suggestions, or examples would be greatly appreciated. If you've faced a similar situation and found a workaround, please share your experience!\n\nThank you in advance for your help!",
      "votes": null
    },
    {
      "id": "2942686",
      "postDate": "08/01/2024 02:22:52",
      "content": "<p><a href=\"https://github.com/Kaggle/kaggle-api\" target=\"_blank\">kaggle has api for that</a></p>",
      "rawMarkdown": "[kaggle has api for that](https://github.com/Kaggle/kaggle-api)",
      "votes": null
    },
    {
      "id": "2943152",
      "postDate": "08/01/2024 11:38:15",
      "content": "<p>You can try preprocessing data in kaggle itself and then you can download zipped outputs.<br>\nsome people have converted dicom files to PNG and created dataset, you can utilize those or try same thing yourself.</p>\n<p>in case you want same file and folder structure for debugging purpose:<br>\nI have created a tiny debug dataset <a href=\"https://www.kaggle.com/datasets/rohitchaudhari25/rsna-lsdc-2024-submission-debug-dataset\" target=\"_blank\">RSNA_LSDC_2024_submission_debug_dataset</a> for debugging purpose.</p>\n<p>source code can be found here <a href=\"https://www.kaggle.com/code/rohitchaudhari25/rsna-lsdc-2024-submission-debug/notebook\" target=\"_blank\">RSNA_LSDC_2024_submission_debug</a></p>\n<p>You can download whole dataset in batches using idea in above notebook with few modifications.<br>\nonly heavy folder is train_images(~28 gb)<br>\nI have chosen to copy few files from train_images folder in <code>kaggle/working/</code> directory.<br>\nyou can use same idea of copying few study_id folders each time from train_images.<br>\nyou can zip them and download  in say batches of 100 study_ids at once. (~20 - batches would suffice)<br>\nother files are lightweight and can be downloaded at once.</p>",
      "rawMarkdown": "You can try preprocessing data in kaggle itself and then you can download zipped outputs.\nsome people have converted dicom files to PNG and created dataset, you can utilize those or try same thing yourself.\n\nin case you want same file and folder structure for debugging purpose:\nI have created a tiny debug dataset [RSNA_LSDC_2024_submission_debug_dataset](https://www.kaggle.com/datasets/rohitchaudhari25/rsna-lsdc-2024-submission-debug-dataset) for debugging purpose.\n\nsource code can be found here [RSNA_LSDC_2024_submission_debug](https://www.kaggle.com/code/rohitchaudhari25/rsna-lsdc-2024-submission-debug/notebook)\n\nYou can download whole dataset in batches using idea in above notebook with few modifications.\nonly heavy folder is train_images(~28 gb)\nI have chosen to copy few files from train_images folder in `kaggle/working/` directory.\nyou can use same idea of copying few study_id folders each time from train_images.\nyou can zip them and download  in say batches of 100 study_ids at once. (~20 - batches would suffice)\nother files are lightweight and can be downloaded at once.",
      "votes": null
    },
    {
      "id": "2943161",
      "postDate": "08/01/2024 11:46:58",
      "content": "<p>Thanks i'll check it out <a href=\"https://www.kaggle.com/rohitchaudhari25\" target=\"_blank\">@rohitchaudhari25</a> </p>",
      "rawMarkdown": "Thanks i'll check it out @rohitchaudhari25",
      "votes": null
    },
    {
      "id": "2943755",
      "postDate": "08/01/2024 20:14:11",
      "content": "<p>Do you find alternative way to download the dataset ?</p>",
      "rawMarkdown": "Do you find alternative way to download the dataset ?",
      "votes": null
    },
    {
      "id": "2943912",
      "postDate": "08/02/2024 01:59:06",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/ravshankutkovin\" target=\"_blank\">@ravshankutkovin</a> <a href=\"https://www.kaggle.com/rohitchaudhari25\" target=\"_blank\">@rohitchaudhari25</a> has suggested the alternative i'll explore that. Aso i tried with different sources like roboflow universe but the size of data was very less there.</p>",
      "rawMarkdown": "Hey @ravshankutkovin @rohitchaudhari25 has suggested the alternative i'll explore that. Aso i tried with different sources like roboflow universe but the size of data was very less there.",
      "votes": null
    },
    {
      "id": "3183075",
      "postDate": "04/20/2025 11:00:51",
      "content": "<p><a href=\"https://www.kaggle.com/gbiamgaurav\" target=\"_blank\">@gbiamgaurav</a>, In a competition, I converted the project from dicom to png because it was too big (you can also use jpeg if you want because it will compress more) and then sent it to google drive using the kaggle api. in your case, if you have the api key, you can download 34 gb of data directly with google colab without changing anything and save it to the drive. this solves the bandwith problem and you can also download it in small chunks later.</p>",
      "rawMarkdown": "gbiamgaurav, In a competition, I converted the project from dicom to png because it was too big (you can also use jpeg if you want because it will compress more) and then sent it to google drive using the kaggle api. in your case, if you have the api key, you can download 34 gb of data directly with google colab without changing anything and save it to the drive. this solves the bandwith problem and you can also download it in small chunks later.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2942686,
      "author_name": "ibinti",
      "author_url": "",
      "post_date": "08/01/2024 02:22:52",
      "content": "<p><a href=\"https://github.com/Kaggle/kaggle-api\" target=\"_blank\">kaggle has api for that</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2943152,
      "author_name": "rohitchaudhari25",
      "author_url": "",
      "post_date": "08/01/2024 11:38:15",
      "content": "<p>You can try preprocessing data in kaggle itself and then you can download zipped outputs.<br>\nsome people have converted dicom files to PNG and created dataset, you can utilize those or try same thing yourself.</p>\n<p>in case you want same file and folder structure for debugging purpose:<br>\nI have created a tiny debug dataset <a href=\"https://www.kaggle.com/datasets/rohitchaudhari25/rsna-lsdc-2024-submission-debug-dataset\" target=\"_blank\">RSNA_LSDC_2024_submission_debug_dataset</a> for debugging purpose.</p>\n<p>source code can be found here <a href=\"https://www.kaggle.com/code/rohitchaudhari25/rsna-lsdc-2024-submission-debug/notebook\" target=\"_blank\">RSNA_LSDC_2024_submission_debug</a></p>\n<p>You can download whole dataset in batches using idea in above notebook with few modifications.<br>\nonly heavy folder is train_images(~28 gb)<br>\nI have chosen to copy few files from train_images folder in <code>kaggle/working/</code> directory.<br>\nyou can use same idea of copying few study_id folders each time from train_images.<br>\nyou can zip them and download  in say batches of 100 study_ids at once. (~20 - batches would suffice)<br>\nother files are lightweight and can be downloaded at once.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2943161,
          "author_name": "gbiamgaurav",
          "author_url": "",
          "post_date": "08/01/2024 11:46:58",
          "content": "<p>Thanks i'll check it out <a href=\"https://www.kaggle.com/rohitchaudhari25\" target=\"_blank\">@rohitchaudhari25</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2943755,
      "author_name": "",
      "author_url": "",
      "post_date": "08/01/2024 20:14:11",
      "content": "<p>Do you find alternative way to download the dataset ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2943912,
          "author_name": "gbiamgaurav",
          "author_url": "",
          "post_date": "08/02/2024 01:59:06",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/ravshankutkovin\" target=\"_blank\">@ravshankutkovin</a> <a href=\"https://www.kaggle.com/rohitchaudhari25\" target=\"_blank\">@rohitchaudhari25</a> has suggested the alternative i'll explore that. Aso i tried with different sources like roboflow universe but the size of data was very less there.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3183075,
      "author_name": "orvile",
      "author_url": "",
      "post_date": "04/20/2025 11:00:51",
      "content": "<p><a href=\"https://www.kaggle.com/gbiamgaurav\" target=\"_blank\">@gbiamgaurav</a>, In a competition, I converted the project from dicom to png because it was too big (you can also use jpeg if you want because it will compress more) and then sent it to google drive using the kaggle api. in your case, if you have the api key, you can download 34 gb of data directly with google colab without changing anything and save it to the drive. this solves the bandwith problem and you can also download it in small chunks later.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2942650": "Hi Folks!\n\nI hope you're all doing well. I'm currently working on this project but the dataset exceeds 34GB in size. Given the substantial size, downloading the entire dataset in one go is not feasible due to time constraints and bandwidth limitations.\n\nI'd like to ask for advice on a couple of points:\n\nIs there an alternative way to download the dataset in smaller chunks (e.g., in MB)? This would help manage the download process more efficiently and avoid potential interruptions or failures during the download.\n\nAre there any tools or scripts that you recommend for this purpose? Perhaps something that can automate the process and resume downloading in case of any disruptions?\n\nAny insights, suggestions, or examples would be greatly appreciated. If you've faced a similar situation and found a workaround, please share your experience!\n\nThank you in advance for your help!",
    "2942686": "[kaggle has api for that](https://github.com/Kaggle/kaggle-api)",
    "2943152": "You can try preprocessing data in kaggle itself and then you can download zipped outputs.\nsome people have converted dicom files to PNG and created dataset, you can utilize those or try same thing yourself.\n\nin case you want same file and folder structure for debugging purpose:\nI have created a tiny debug dataset [RSNA_LSDC_2024_submission_debug_dataset](https://www.kaggle.com/datasets/rohitchaudhari25/rsna-lsdc-2024-submission-debug-dataset) for debugging purpose.\n\nsource code can be found here [RSNA_LSDC_2024_submission_debug](https://www.kaggle.com/code/rohitchaudhari25/rsna-lsdc-2024-submission-debug/notebook)\n\nYou can download whole dataset in batches using idea in above notebook with few modifications.\nonly heavy folder is train_images(~28 gb)\nI have chosen to copy few files from train_images folder in `kaggle/working/` directory.\nyou can use same idea of copying few study_id folders each time from train_images.\nyou can zip them and download  in say batches of 100 study_ids at once. (~20 - batches would suffice)\nother files are lightweight and can be downloaded at once.",
    "2943161": "Thanks i'll check it out @rohitchaudhari25",
    "2943755": "Do you find alternative way to download the dataset ?",
    "2943912": "Hey @ravshankutkovin @rohitchaudhari25 has suggested the alternative i'll explore that. Aso i tried with different sources like roboflow universe but the size of data was very less there.",
    "3183075": "gbiamgaurav, In a competition, I converted the project from dicom to png because it was too big (you can also use jpeg if you want because it will compress more) and then sent it to google drive using the kaggle api. in your case, if you have the api key, you can download 34 gb of data directly with google colab without changing anything and save it to the drive. this solves the bandwith problem and you can also download it in small chunks later."
  },
  "source": "meta"
}