{
  "id": 240411,
  "title": "Missing studies from Kaggle API download",
  "url": "/competitions/siim-covid19-detection/discussion/240411",
  "author_name": "Ian Pan",
  "post_date": "2021-05-19T17:35:58.679000",
  "votes": 23,
  "comment_count": 19,
  "views": 0,
  "content": "<p>I downloaded the competition data via the Kaggle API. It seems there are some missing studies. I only have 5,819/6,054 train studies and 1,178/1,214 test studies. Are other people having the same issue?</p>",
  "messages": [
    {
      "id": 1315297,
      "postDate": "2021-05-19T17:35:58.680Z",
      "content": "<p>I downloaded the competition data via the Kaggle API. It seems there are some missing studies. I only have 5,819/6,054 train studies and 1,178/1,214 test studies. Are other people having the same issue?</p>",
      "rawMarkdown": "I downloaded the competition data via the Kaggle API. It seems there are some missing studies. I only have 5,819/6,054 train studies and 1,178/1,214 test studies. Are other people having the same issue?",
      "votes": 23
    },
    {
      "id": 1318222,
      "postDate": "2021-05-22T05:58:34.287Z",
      "content": "<p>Final update: this has been fixed and verified. All bundles are now containing the correct full number of files.</p>",
      "rawMarkdown": "Final update: this has been fixed and verified. All bundles are now containing the correct full number of files.",
      "votes": 3,
      "replies": [
        {
          "id": 1318275,
          "postDate": "2021-05-22T07:03:24.557Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1315546,
      "postDate": "2021-05-19T22:36:10.893Z",
      "content": "<p>Update: We've replicated the error and are rebuilding the bundle. It appears there was some corruption that occurred in its .zip packaging. We expect the issue to be resolved in the next 24 hours. Sorry for the inconvenience, and we appreciate your patience as we correct it.</p>\n<p>In the meantime, the bundle available in Kaggle notebooks with the dataset loaded is in tact and working, so that is an alternative. </p>",
      "rawMarkdown": "Update: We've replicated the error and are rebuilding the bundle. It appears there was some corruption that occurred in its .zip packaging. We expect the issue to be resolved in the next 24 hours. Sorry for the inconvenience, and we appreciate your patience as we correct it.\n\nIn the meantime, the bundle available in Kaggle notebooks with the dataset loaded is in tact and working, so that is an alternative. ",
      "votes": 2
    },
    {
      "id": 1315318,
      "postDate": "2021-05-19T17:58:54.093Z",
      "content": "<p>I've got the same problem. I got missing files through kaggle notebook.</p>",
      "rawMarkdown": "I've got the same problem. I got missing files through kaggle notebook.",
      "votes": 2,
      "replies": [
        {
          "id": 1315483,
          "postDate": "2021-05-19T20:41:45.570Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/osciiart\" target=\"_blank\">@osciiart</a>! Which files are you missing? What are the counts?</p>",
          "rawMarkdown": "Thanks @osciiart! Which files are you missing? What are the counts?"
        }
      ]
    },
    {
      "id": 1325751,
      "postDate": "2021-05-28T01:38:27.177Z",
      "content": "<p>maybe try to uninstall and reinstall the kaggle api with the latest version. I had a similar issue in an another competition few weeks ago, it was related to an outdated kaggle api.</p>",
      "rawMarkdown": "maybe try to uninstall and reinstall the kaggle api with the latest version. I had a similar issue in an another competition few weeks ago, it was related to an outdated kaggle api.",
      "replies": [
        {
          "id": 1326887,
          "postDate": "2021-05-28T17:58:42.203Z",
          "content": "<p>hello thank you this worked!</p>",
          "rawMarkdown": "hello thank you this worked!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1325539,
      "postDate": "2021-05-27T20:18:12.703Z",
      "content": "<p>I am having the same exact issue I only got train_study_level.csv, train_image_level.csv, sample submission and in addition I only got like 30 images and im not sure if it is from train or test</p>",
      "rawMarkdown": "I am having the same exact issue I only got train_study_level.csv, train_image_level.csv, sample submission and in addition I only got like 30 images and im not sure if it is from train or test",
      "replies": [
        {
          "id": 1325571,
          "postDate": "2021-05-27T20:47:48.620Z",
          "content": "<p>How large was the file that downloaded for you? Should have been a single zip.</p>",
          "rawMarkdown": "How large was the file that downloaded for you? Should have been a single zip."
        },
        {
          "id": 1325577,
          "postDate": "2021-05-27T20:50:56.823Z",
          "content": "<p>Im not sure the size of the file, because it did not say when I downloaded it i got multiple zip files not a single one</p>",
          "rawMarkdown": "Im not sure the size of the file, because it did not say when I downloaded it i got multiple zip files not a single one"
        },
        {
          "id": 1325588,
          "postDate": "2021-05-27T20:59:24.580Z",
          "content": "<p>Ahh, okay. Were you intending to download separate files, or would you like the full archive? That should be <code>kaggle competitions download -c siim-covid19-detection</code></p>\n<p>I'll look into why the files-level download might not be downloading everything.</p>",
          "rawMarkdown": "Ahh, okay. Were you intending to download separate files, or would you like the full archive? That should be `kaggle competitions download -c siim-covid19-detection`\n\nI'll look into why the files-level download might not be downloading everything."
        },
        {
          "id": 1325623,
          "postDate": "2021-05-27T21:24:09.340Z",
          "content": "<p>I put 'kaggle competitions download -c siim-covid19-detection' in colab and got seperate files but i do want the full archive</p>",
          "rawMarkdown": "I put 'kaggle competitions download -c siim-covid19-detection' in colab and got seperate files but i do want the full archive"
        },
        {
          "id": 1325631,
          "postDate": "2021-05-27T21:43:19.320Z",
          "content": "<p>Colab might have a (very) old version of the Kaggle API. Can you try updating it? (I don't recall if <code>pip</code> works on Colab, but <code>pip install kaggle --upgrade</code> should do it.) The command you mention <em>should</em> download the full archive.</p>",
          "rawMarkdown": "Colab might have a (very) old version of the Kaggle API. Can you try updating it? (I don't recall if `pip` works on Colab, but `pip install kaggle --upgrade` should do it.) The command you mention *should* download the full archive."
        },
        {
          "id": 1326892,
          "postDate": "2021-05-28T17:59:19.407Z",
          "content": "<p>I uninstalled and reinstalled kaggle and it worked but thank you for your help</p>",
          "rawMarkdown": "I uninstalled and reinstalled kaggle and it worked but thank you for your help"
        }
      ]
    },
    {
      "id": 1323478,
      "postDate": "2021-05-26T09:00:56.043Z",
      "content": "<p>Does anyone know of a way to download only the missing studies instead of re-downloading the entire 84GB archive? Unfortunately, on my connection, it takes around 8 hours just to download the archive. I've tried:</p>\n<p><code>kaggle competitions download -f &lt;study_path&gt; -c siim-covid19-detection</code></p>\n<p>but I get 404 - Not Found errors. I do understand that -f is meant to work with files and not directories but I don't see any other options in the <a href=\"https://github.com/Kaggle/kaggle-api\" target=\"_blank\">API documentation</a>.</p>\n<p>Besides, <code>kaggle competitions files -c siim-covid19-detection</code> only returns a handful of files. So, I'm not sure I'll be able to download even if I specify the full file paths. Has anyone tried this?</p>\n<p>Thanks in advance!</p>",
      "rawMarkdown": "Does anyone know of a way to download only the missing studies instead of re-downloading the entire 84GB archive? Unfortunately, on my connection, it takes around 8 hours just to download the archive. I've tried:\n\n`kaggle competitions download -f <study_path> -c siim-covid19-detection`\n\nbut I get 404 - Not Found errors. I do understand that -f is meant to work with files and not directories but I don't see any other options in the [API documentation](https://github.com/Kaggle/kaggle-api).\n\nBesides, `kaggle competitions files -c siim-covid19-detection` only returns a handful of files. So, I'm not sure I'll be able to download even if I specify the full file paths. Has anyone tried this?\n\nThanks in advance!",
      "replies": [
        {
          "id": 1325548,
          "postDate": "2021-05-27T20:33:17.893Z",
          "content": "<p>I did 'kaggle competitions files -c siim-covid19-detection' aswell and I only got some of the items, have you found a way to get all of them?</p>",
          "rawMarkdown": "I did 'kaggle competitions files -c siim-covid19-detection' aswell and I only got some of the items, have you found a way to get all of them?"
        },
        {
          "id": 1326021,
          "postDate": "2021-05-28T07:01:41.023Z",
          "content": "<p>I ended up re-downloading the archive overnight.</p>\n<p>In the new archive, 201 files had bad CRCs and a visual inspection revealed that the files were, indeed, corrupt. Since no one else has reported this issue, I assume, the corruption must have occurred at my end.</p>\n<p>Fortunately, even though the files were corrupt, at least now I had full file paths to use with <code>kaggle competitions download -f &lt;dcm_path&gt; -c siim-covid19-detection</code>. After running into a \"Too many Requests\" error after approx. 150 individual downloads yesterday, I've just finished downloading the remaining files.</p>\n<p>I now have 6334 &amp; 1263 DCMs in the train and test directories, respectively.</p>",
          "rawMarkdown": "I ended up re-downloading the archive overnight.\n\nIn the new archive, 201 files had bad CRCs and a visual inspection revealed that the files were, indeed, corrupt. Since no one else has reported this issue, I assume, the corruption must have occurred at my end.\n\nFortunately, even though the files were corrupt, at least now I had full file paths to use with `kaggle competitions download -f <dcm_path> -c siim-covid19-detection`. After running into a \"Too many Requests\" error after approx. 150 individual downloads yesterday, I've just finished downloading the remaining files.\n\nI now have 6334 & 1263 DCMs in the train and test directories, respectively."
        }
      ]
    },
    {
      "id": 1315477,
      "postDate": "2021-05-19T20:30:41.337Z",
      "content": "<p>Thanks for the report. We are investigating this!</p>",
      "rawMarkdown": "Thanks for the report. We are investigating this!"
    },
    {
      "id": 1315338,
      "postDate": "2021-05-19T18:10:20.993Z",
      "content": "<p>Hi, I simply checked my cases.</p>\n<pre><code>len(glob.glob('../input/siim-covid19-detection/train/*'))\n\nlen(glob.glob('../input/siim-covid19-detection/test/*'))\n</code></pre>\n<table>\n<thead>\n<tr>\n<th>In Kaggle</th>\n<th>Kaggle API</th>\n<th>Etc. <a href=\"https://www.kaggle.com/xhlulu/siim-covid19-resized-to-256px-jpg\" target=\"_blank\">resized-to-256px-jpg</a></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>train / test</td>\n<td>train / test</td>\n<td>train / test</td>\n</tr>\n<tr>\n<td>6,054 / 1,214</td>\n<td>5,819 / 1,178</td>\n<td>6,334 / 1,263</td>\n</tr>\n</tbody>\n</table>\n<p>There is same issue for me(maybe I'm wrong).<br>\nI hope it helps you.</p>\n<p>p.s. It seems the main cause of my <code>submission score error</code> - <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/240332\" target=\"_blank\">link</a>.<br>\n😅</p>",
      "rawMarkdown": "Hi, I simply checked my cases.\n```\nlen(glob.glob('../input/siim-covid19-detection/train/*'))\n\nlen(glob.glob('../input/siim-covid19-detection/test/*'))\n```\n\n| In Kaggle     | Kaggle API| Etc. [resized-to-256px-jpg](https://www.kaggle.com/xhlulu/siim-covid19-resized-to-256px-jpg) | \n| --- | --- |--- |\n|train / test | train / test | train / test |\n|6,054 / 1,214 | 5,819 / 1,178 | 6,334 / 1,263 |\n\nThere is same issue for me(maybe I'm wrong).\nI hope it helps you.\n\np.s. It seems the main cause of my `submission score error` - [link](https://www.kaggle.com/c/siim-covid19-detection/discussion/240332).\n😅\n"
    }
  ],
  "comments": [
    {
      "id": 1318222,
      "author_name": "Julia Elliott",
      "author_url": "",
      "post_date": "2021-05-22T05:58:34.287000",
      "content": "<p>Final update: this has been fixed and verified. All bundles are now containing the correct full number of files.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1318275,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-22T07:03:24.557000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1315546,
      "author_name": "Julia Elliott",
      "author_url": "",
      "post_date": "2021-05-19T22:36:10.893000",
      "content": "<p>Update: We've replicated the error and are rebuilding the bundle. It appears there was some corruption that occurred in its .zip packaging. We expect the issue to be resolved in the next 24 hours. Sorry for the inconvenience, and we appreciate your patience as we correct it.</p>\n<p>In the meantime, the bundle available in Kaggle notebooks with the dataset loaded is in tact and working, so that is an alternative. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1315318,
      "author_name": "OsciiArt",
      "author_url": "",
      "post_date": "2021-05-19T17:58:54.093000",
      "content": "<p>I've got the same problem. I got missing files through kaggle notebook.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1315483,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2021-05-19T20:41:45.570000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/osciiart\" target=\"_blank\">@osciiart</a>! Which files are you missing? What are the counts?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1325751,
      "author_name": "Shiro",
      "author_url": "",
      "post_date": "2021-05-28T01:38:27.177000",
      "content": "<p>maybe try to uninstall and reinstall the kaggle api with the latest version. I had a similar issue in an another competition few weeks ago, it was related to an outdated kaggle api.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1326887,
          "author_name": "Simonsob",
          "author_url": "",
          "post_date": "2021-05-28T17:58:42.203000",
          "content": "<p>hello thank you this worked!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1325539,
      "author_name": "Simonsob",
      "author_url": "",
      "post_date": "2021-05-27T20:18:12.703000",
      "content": "<p>I am having the same exact issue I only got train_study_level.csv, train_image_level.csv, sample submission and in addition I only got like 30 images and im not sure if it is from train or test</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1325571,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2021-05-27T20:47:48.620000",
          "content": "<p>How large was the file that downloaded for you? Should have been a single zip.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1325577,
          "author_name": "Simonsob",
          "author_url": "",
          "post_date": "2021-05-27T20:50:56.823000",
          "content": "<p>Im not sure the size of the file, because it did not say when I downloaded it i got multiple zip files not a single one</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1325588,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2021-05-27T20:59:24.580000",
          "content": "<p>Ahh, okay. Were you intending to download separate files, or would you like the full archive? That should be <code>kaggle competitions download -c siim-covid19-detection</code></p>\n<p>I'll look into why the files-level download might not be downloading everything.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1325623,
          "author_name": "Simonsob",
          "author_url": "",
          "post_date": "2021-05-27T21:24:09.340000",
          "content": "<p>I put 'kaggle competitions download -c siim-covid19-detection' in colab and got seperate files but i do want the full archive</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1325631,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2021-05-27T21:43:19.320000",
          "content": "<p>Colab might have a (very) old version of the Kaggle API. Can you try updating it? (I don't recall if <code>pip</code> works on Colab, but <code>pip install kaggle --upgrade</code> should do it.) The command you mention <em>should</em> download the full archive.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1326892,
          "author_name": "Simonsob",
          "author_url": "",
          "post_date": "2021-05-28T17:59:19.407000",
          "content": "<p>I uninstalled and reinstalled kaggle and it worked but thank you for your help</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1323478,
      "author_name": "akki",
      "author_url": "",
      "post_date": "2021-05-26T09:00:56.043000",
      "content": "<p>Does anyone know of a way to download only the missing studies instead of re-downloading the entire 84GB archive? Unfortunately, on my connection, it takes around 8 hours just to download the archive. I've tried:</p>\n<p><code>kaggle competitions download -f &lt;study_path&gt; -c siim-covid19-detection</code></p>\n<p>but I get 404 - Not Found errors. I do understand that -f is meant to work with files and not directories but I don't see any other options in the <a href=\"https://github.com/Kaggle/kaggle-api\" target=\"_blank\">API documentation</a>.</p>\n<p>Besides, <code>kaggle competitions files -c siim-covid19-detection</code> only returns a handful of files. So, I'm not sure I'll be able to download even if I specify the full file paths. Has anyone tried this?</p>\n<p>Thanks in advance!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1325548,
          "author_name": "Simonsob",
          "author_url": "",
          "post_date": "2021-05-27T20:33:17.893000",
          "content": "<p>I did 'kaggle competitions files -c siim-covid19-detection' aswell and I only got some of the items, have you found a way to get all of them?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1326021,
          "author_name": "akki",
          "author_url": "",
          "post_date": "2021-05-28T07:01:41.023000",
          "content": "<p>I ended up re-downloading the archive overnight.</p>\n<p>In the new archive, 201 files had bad CRCs and a visual inspection revealed that the files were, indeed, corrupt. Since no one else has reported this issue, I assume, the corruption must have occurred at my end.</p>\n<p>Fortunately, even though the files were corrupt, at least now I had full file paths to use with <code>kaggle competitions download -f &lt;dcm_path&gt; -c siim-covid19-detection</code>. After running into a \"Too many Requests\" error after approx. 150 individual downloads yesterday, I've just finished downloading the remaining files.</p>\n<p>I now have 6334 &amp; 1263 DCMs in the train and test directories, respectively.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1315477,
      "author_name": "Julia Elliott",
      "author_url": "",
      "post_date": "2021-05-19T20:30:41.337000",
      "content": "<p>Thanks for the report. We are investigating this!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1315338,
      "author_name": "Heroseo",
      "author_url": "",
      "post_date": "2021-05-19T18:10:20.993000",
      "content": "<p>Hi, I simply checked my cases.</p>\n<pre><code>len(glob.glob('../input/siim-covid19-detection/train/*'))\n\nlen(glob.glob('../input/siim-covid19-detection/test/*'))\n</code></pre>\n<table>\n<thead>\n<tr>\n<th>In Kaggle</th>\n<th>Kaggle API</th>\n<th>Etc. <a href=\"https://www.kaggle.com/xhlulu/siim-covid19-resized-to-256px-jpg\" target=\"_blank\">resized-to-256px-jpg</a></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>train / test</td>\n<td>train / test</td>\n<td>train / test</td>\n</tr>\n<tr>\n<td>6,054 / 1,214</td>\n<td>5,819 / 1,178</td>\n<td>6,334 / 1,263</td>\n</tr>\n</tbody>\n</table>\n<p>There is same issue for me(maybe I'm wrong).<br>\nI hope it helps you.</p>\n<p>p.s. It seems the main cause of my <code>submission score error</code> - <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/240332\" target=\"_blank\">link</a>.<br>\n😅</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1315297": "I downloaded the competition data via the Kaggle API. It seems there are some missing studies. I only have 5,819/6,054 train studies and 1,178/1,214 test studies. Are other people having the same issue?",
    "1318222": "Final update: this has been fixed and verified. All bundles are now containing the correct full number of files.",
    "1315546": "Update: We've replicated the error and are rebuilding the bundle. It appears there was some corruption that occurred in its .zip packaging. We expect the issue to be resolved in the next 24 hours. Sorry for the inconvenience, and we appreciate your patience as we correct it.\n\nIn the meantime, the bundle available in Kaggle notebooks with the dataset loaded is in tact and working, so that is an alternative. ",
    "1315318": "I've got the same problem. I got missing files through kaggle notebook.",
    "1325751": "maybe try to uninstall and reinstall the kaggle api with the latest version. I had a similar issue in an another competition few weeks ago, it was related to an outdated kaggle api.",
    "1325539": "I am having the same exact issue I only got train_study_level.csv, train_image_level.csv, sample submission and in addition I only got like 30 images and im not sure if it is from train or test",
    "1323478": "Does anyone know of a way to download only the missing studies instead of re-downloading the entire 84GB archive? Unfortunately, on my connection, it takes around 8 hours just to download the archive. I've tried:\n\n`kaggle competitions download -f <study_path> -c siim-covid19-detection`\n\nbut I get 404 - Not Found errors. I do understand that -f is meant to work with files and not directories but I don't see any other options in the [API documentation](https://github.com/Kaggle/kaggle-api).\n\nBesides, `kaggle competitions files -c siim-covid19-detection` only returns a handful of files. So, I'm not sure I'll be able to download even if I specify the full file paths. Has anyone tried this?\n\nThanks in advance!",
    "1315477": "Thanks for the report. We are investigating this!",
    "1315338": "Hi, I simply checked my cases.\n```\nlen(glob.glob('../input/siim-covid19-detection/train/*'))\n\nlen(glob.glob('../input/siim-covid19-detection/test/*'))\n```\n\n| In Kaggle     | Kaggle API| Etc. [resized-to-256px-jpg](https://www.kaggle.com/xhlulu/siim-covid19-resized-to-256px-jpg) | \n| --- | --- |--- |\n|train / test | train / test | train / test |\n|6,054 / 1,214 | 5,819 / 1,178 | 6,334 / 1,263 |\n\nThere is same issue for me(maybe I'm wrong).\nI hope it helps you.\n\np.s. It seems the main cause of my `submission score error` - [link](https://www.kaggle.com/c/siim-covid19-detection/discussion/240332).\n😅\n"
  }
}