{
  "id": 198440,
  "title": "Register Large NB Output Files as Datasets without Re-upload",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/198440",
  "author_name": "Sreevishnu Damodaran",
  "post_date": "2020-11-21T08:43:43.041000",
  "votes": 14,
  "comment_count": 0,
  "views": 0,
  "content": "<p><strong>Dear Kagglers,</strong></p>\n<p>A little insight for everyone looking for ways to Upload and Register the outputs of notebook files as Kaggle Datasets without re-uploading it.</p>\n<p>Downloading the output files and uploading it again to the Datasets page can be particularly difficult when the datasets size is huge and also if the number of files are large. This is often the case with Computer Vision Competitions and Projects.</p>\n<p>The notebooks also runs into a lot of problems when a 'Save Version' or 'Commit' is done with the scenario of large data size and number mentioned above. It is particularly difficult to 'Save Version' when there are more than 50 files in as notebook outputs or when the size exceed several GBs.</p>\n<p>I came across this issue while trying to register a modified dataset for the HubMap competition which had 26k+ output files.</p>\n<p><strong>Notebook and Dataset</strong></p>\n<p><a href=\"https://www.kaggle.com/sreevishnudamodaran/build-custom-coco-annotations-512x512-tiled\" target=\"_blank\">https://www.kaggle.com/sreevishnudamodaran/build-custom-coco-annotations-512x512-tiled</a></p>\n<p><a href=\"https://www.kaggle.com/sreevishnudamodaran/hubmap-coco-dataset-512x512-tiled\" target=\"_blank\">https://www.kaggle.com/sreevishnudamodaran/hubmap-coco-dataset-512x512-tiled</a></p>\n<h4>Steps:</h4>\n<p><strong>1. Organize all the files to a proper folder structure that is to be registered.</strong><br>\nEx: </p>\n<pre><code>    - images\n          - image_1\n          - image_2\n                  .\n                  .\n     - masks\n           - mask_1\n           - mask_2\n                  .\n                  .\n</code></pre>\n<p>or</p>\n<pre><code>    - coco_train\n        - images\n            - image_1\n            - image_2\n                   .\n                   .\n        - train.json (All annotations in coco format with proper relative path of images)\n</code></pre>\n<p><strong>2.  Create a Zip archive. Kaggle will automatically extract and unpack on uploading and registering the zip.</strong></p>\n<pre><code>!zip -r dataset_name.zip ./coco_train\n</code></pre>\n<ul>\n<li>Remove the unarchived folder if needed</li>\n</ul>\n<pre><code>!rm -R coco_train\n</code></pre>\n<p><strong>3.  Check the new file size in MBs</strong></p>\n<pre><code>!ls -ahl\n</code></pre>\n<p><strong>4. Click 'Save Version' to save the notebook and in the versions section, click 'Go to viewer' .</strong></p>\n<p><strong>5. Go to the bottom and click 'New Dataset'. This may be a little tricky to find so here's a screenshot.</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5532218%2F1567943cce91cff8fc4e305fc26c9c33%2Fstep.png?generation=1605948034551473&amp;alt=media\" alt=\"\"></p>\n<p><strong>7. Add details, save and then go the Data section to see your dataset.</strong></p>\n<p><strong>8.Done. Add further details such as subtitle, metadata, tags, license etc as needed from the datasets page.</strong></p>\n<h2>UPDATE:</h2>\n<h2>Alternate Method using Kaggle API directly from Notebook</h2>\n<p><strong>1. Organize into data into proper folders as specified in the step 1 above</strong></p>\n<p><strong>2. Run the below cell with the right kaggle_username &amp; kaggle_api_key. Get the kaggle_api_key from the 'Account' tab of your user profile (https://www.kaggle.com//account) and select 'Create API Token'. This will trigger the download of kaggle.json, a file containing your API credentials.</strong></p>\n<pre><code>%%bash\ntouch /root/.kaggle/kaggle.json\necho '{\"username\":\"&lt;kaggle_username&gt;\",\"key\":\"&lt;kaggle_api_key&gt;\"}' &gt; /root/.kaggle/kaggle.json\ncat /root/.kaggle/kaggle.json\nchmod 600 /root/.kaggle/kaggle.json\n</code></pre>\n<p><strong>3. Test if the API works by listing datasets.</strong></p>\n<pre><code>!kaggle datasets list -s \"&lt;search_word&gt;\"\n</code></pre>\n<p><strong>4. Generate the dataset-metadata.json</strong></p>\n<pre><code>!kaggle datasets init -p &lt;dataset_folder_name&gt;\n</code></pre>\n<p><strong>5. Edit the metadata in the file.</strong></p>\n<pre><code>import json\nf = open('./&lt;dataset_folder_name&gt;/dataset-metadata.json')\nmeta = json.load(f)\n\nmeta['title'] = \"&lt;dataset_title&gt;\"\nmeta['id'] = \"&lt;kaggle_username&gt;/&lt;dataset_name&gt;\"\n\nprint(meta)\n\nwith open('&lt;dataset_folder_name&gt;/dataset-metadata.json', 'w') as json_file:\n  json.dump(meta, json_file, indent=4)\n</code></pre>\n<p><strong>6. Start the Dataset creation and upload.</strong></p>\n<pre><code>!kaggle datasets create -p ./&lt;dataset_folder_name&gt;/ -r zip\n</code></pre>\n<p>Note:</p>\n<ul>\n<li>Specify 'zip' mode for directory upload with compression. Needed only if there are sub-directories within the dataset directory</li>\n<li>It might take a few minutes to see an output.</li>\n</ul>\n<h4>Hope everyone finds this useful.</h4>",
  "messages": [
    {
      "id": 1085858,
      "postDate": "2020-11-21T08:43:43.040Z",
      "content": "<p><strong>Dear Kagglers,</strong></p>\n<p>A little insight for everyone looking for ways to Upload and Register the outputs of notebook files as Kaggle Datasets without re-uploading it.</p>\n<p>Downloading the output files and uploading it again to the Datasets page can be particularly difficult when the datasets size is huge and also if the number of files are large. This is often the case with Computer Vision Competitions and Projects.</p>\n<p>The notebooks also runs into a lot of problems when a 'Save Version' or 'Commit' is done with the scenario of large data size and number mentioned above. It is particularly difficult to 'Save Version' when there are more than 50 files in as notebook outputs or when the size exceed several GBs.</p>\n<p>I came across this issue while trying to register a modified dataset for the HubMap competition which had 26k+ output files.</p>\n<p><strong>Notebook and Dataset</strong></p>\n<p><a href=\"https://www.kaggle.com/sreevishnudamodaran/build-custom-coco-annotations-512x512-tiled\" target=\"_blank\">https://www.kaggle.com/sreevishnudamodaran/build-custom-coco-annotations-512x512-tiled</a></p>\n<p><a href=\"https://www.kaggle.com/sreevishnudamodaran/hubmap-coco-dataset-512x512-tiled\" target=\"_blank\">https://www.kaggle.com/sreevishnudamodaran/hubmap-coco-dataset-512x512-tiled</a></p>\n<h4>Steps:</h4>\n<p><strong>1. Organize all the files to a proper folder structure that is to be registered.</strong><br>\nEx: </p>\n<pre><code>    - images\n          - image_1\n          - image_2\n                  .\n                  .\n     - masks\n           - mask_1\n           - mask_2\n                  .\n                  .\n</code></pre>\n<p>or</p>\n<pre><code>    - coco_train\n        - images\n            - image_1\n            - image_2\n                   .\n                   .\n        - train.json (All annotations in coco format with proper relative path of images)\n</code></pre>\n<p><strong>2.  Create a Zip archive. Kaggle will automatically extract and unpack on uploading and registering the zip.</strong></p>\n<pre><code>!zip -r dataset_name.zip ./coco_train\n</code></pre>\n<ul>\n<li>Remove the unarchived folder if needed</li>\n</ul>\n<pre><code>!rm -R coco_train\n</code></pre>\n<p><strong>3.  Check the new file size in MBs</strong></p>\n<pre><code>!ls -ahl\n</code></pre>\n<p><strong>4. Click 'Save Version' to save the notebook and in the versions section, click 'Go to viewer' .</strong></p>\n<p><strong>5. Go to the bottom and click 'New Dataset'. This may be a little tricky to find so here's a screenshot.</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5532218%2F1567943cce91cff8fc4e305fc26c9c33%2Fstep.png?generation=1605948034551473&amp;alt=media\" alt=\"\"></p>\n<p><strong>7. Add details, save and then go the Data section to see your dataset.</strong></p>\n<p><strong>8.Done. Add further details such as subtitle, metadata, tags, license etc as needed from the datasets page.</strong></p>\n<h2>UPDATE:</h2>\n<h2>Alternate Method using Kaggle API directly from Notebook</h2>\n<p><strong>1. Organize into data into proper folders as specified in the step 1 above</strong></p>\n<p><strong>2. Run the below cell with the right kaggle_username &amp; kaggle_api_key. Get the kaggle_api_key from the 'Account' tab of your user profile (https://www.kaggle.com//account) and select 'Create API Token'. This will trigger the download of kaggle.json, a file containing your API credentials.</strong></p>\n<pre><code>%%bash\ntouch /root/.kaggle/kaggle.json\necho '{\"username\":\"&lt;kaggle_username&gt;\",\"key\":\"&lt;kaggle_api_key&gt;\"}' &gt; /root/.kaggle/kaggle.json\ncat /root/.kaggle/kaggle.json\nchmod 600 /root/.kaggle/kaggle.json\n</code></pre>\n<p><strong>3. Test if the API works by listing datasets.</strong></p>\n<pre><code>!kaggle datasets list -s \"&lt;search_word&gt;\"\n</code></pre>\n<p><strong>4. Generate the dataset-metadata.json</strong></p>\n<pre><code>!kaggle datasets init -p &lt;dataset_folder_name&gt;\n</code></pre>\n<p><strong>5. Edit the metadata in the file.</strong></p>\n<pre><code>import json\nf = open('./&lt;dataset_folder_name&gt;/dataset-metadata.json')\nmeta = json.load(f)\n\nmeta['title'] = \"&lt;dataset_title&gt;\"\nmeta['id'] = \"&lt;kaggle_username&gt;/&lt;dataset_name&gt;\"\n\nprint(meta)\n\nwith open('&lt;dataset_folder_name&gt;/dataset-metadata.json', 'w') as json_file:\n  json.dump(meta, json_file, indent=4)\n</code></pre>\n<p><strong>6. Start the Dataset creation and upload.</strong></p>\n<pre><code>!kaggle datasets create -p ./&lt;dataset_folder_name&gt;/ -r zip\n</code></pre>\n<p>Note:</p>\n<ul>\n<li>Specify 'zip' mode for directory upload with compression. Needed only if there are sub-directories within the dataset directory</li>\n<li>It might take a few minutes to see an output.</li>\n</ul>\n<h4>Hope everyone finds this useful.</h4>",
      "rawMarkdown": "**Dear Kagglers,**\n\nA little insight for everyone looking for ways to Upload and Register the outputs of notebook files as Kaggle Datasets without re-uploading it.\n\nDownloading the output files and uploading it again to the Datasets page can be particularly difficult when the datasets size is huge and also if the number of files are large. This is often the case with Computer Vision Competitions and Projects.\n\nThe notebooks also runs into a lot of problems when a 'Save Version' or 'Commit' is done with the scenario of large data size and number mentioned above. It is particularly difficult to 'Save Version' when there are more than 50 files in as notebook outputs or when the size exceed several GBs.\n\nI came across this issue while trying to register a modified dataset for the HubMap competition which had 26k+ output files.\n\n**Notebook and Dataset**\n\nhttps://www.kaggle.com/sreevishnudamodaran/build-custom-coco-annotations-512x512-tiled\n\nhttps://www.kaggle.com/sreevishnudamodaran/hubmap-coco-dataset-512x512-tiled\n\n\n#### Steps:\n\n**1. Organize all the files to a proper folder structure that is to be registered.**\nEx: \n```\n\t- images\n\t\t  - image_1\n\t\t  - image_2\n\t\t\t\t  .\n\t\t\t\t  .\n\t - masks\n\t\t   - mask_1\n\t\t   - mask_2\n\t\t\t\t  .\n\t\t\t\t  .\n```\nor\n\n```\n\t- coco_train\n\t\t- images\n\t\t\t- image_1\n\t\t\t- image_2\n\t\t\t\t   .\n\t\t\t\t   .\n\t\t- train.json (All annotations in coco format with proper relative path of images)\n```\n\t\t\n**2.  Create a Zip archive. Kaggle will automatically extract and unpack on uploading and registering the zip.**\n \n```python\n!zip -r dataset_name.zip ./coco_train\n```\n\n- Remove the unarchived folder if needed\n\n```python\n!rm -R coco_train\n```\n\n**3.  Check the new file size in MBs**\n\n```python\n!ls -ahl\n```\n\n**4. Click 'Save Version' to save the notebook and in the versions section, click 'Go to viewer' .**\n\n**5. Go to the bottom and click 'New Dataset'. This may be a little tricky to find so here's a screenshot.**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5532218%2F1567943cce91cff8fc4e305fc26c9c33%2Fstep.png?generation=1605948034551473&alt=media)\n\n\n\n\n\n\n**7. Add details, save and then go the Data section to see your dataset.**\n\n**8.Done. Add further details such as subtitle, metadata, tags, license etc as needed from the datasets page.**\n\n\n\n\n\n\n## UPDATE:\n## Alternate Method using Kaggle API directly from Notebook\n\n**1. Organize into data into proper folders as specified in the step 1 above**\n\n**2. Run the below cell with the right kaggle_username & kaggle_api_key. Get the kaggle_api_key from the 'Account' tab of your user profile (https://www.kaggle.com/<username>/account) and select 'Create API Token'. This will trigger the download of kaggle.json, a file containing your API credentials.**\n \n```\n%%bash\ntouch /root/.kaggle/kaggle.json\necho '{\"username\":\"<kaggle_username>\",\"key\":\"<kaggle_api_key>\"}' > /root/.kaggle/kaggle.json\ncat /root/.kaggle/kaggle.json\nchmod 600 /root/.kaggle/kaggle.json\n```\n\n**3. Test if the API works by listing datasets.**\n\n```\n!kaggle datasets list -s \"<search_word>\"\n```\n\n**4. Generate the dataset-metadata.json**\n\n```\n!kaggle datasets init -p <dataset_folder_name>\n```\n\n**5. Edit the metadata in the file.**\n\n```\nimport json\nf = open('./<dataset_folder_name>/dataset-metadata.json')\nmeta = json.load(f)\n\nmeta['title'] = \"<dataset_title>\"\nmeta['id'] = \"<kaggle_username>/<dataset_name>\"\n\nprint(meta)\n\nwith open('<dataset_folder_name>/dataset-metadata.json', 'w') as json_file:\n  json.dump(meta, json_file, indent=4)\n  \n```\n\n\n**6. Start the Dataset creation and upload.**\n\n```\n!kaggle datasets create -p ./<dataset_folder_name>/ -r zip\n```\n\nNote:\n- Specify 'zip' mode for directory upload with compression. Needed only if there are sub-directories within the dataset directory\n- It might take a few minutes to see an output.\n\n#### Hope everyone finds this useful.",
      "votes": 14
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1085858": "**Dear Kagglers,**\n\nA little insight for everyone looking for ways to Upload and Register the outputs of notebook files as Kaggle Datasets without re-uploading it.\n\nDownloading the output files and uploading it again to the Datasets page can be particularly difficult when the datasets size is huge and also if the number of files are large. This is often the case with Computer Vision Competitions and Projects.\n\nThe notebooks also runs into a lot of problems when a 'Save Version' or 'Commit' is done with the scenario of large data size and number mentioned above. It is particularly difficult to 'Save Version' when there are more than 50 files in as notebook outputs or when the size exceed several GBs.\n\nI came across this issue while trying to register a modified dataset for the HubMap competition which had 26k+ output files.\n\n**Notebook and Dataset**\n\nhttps://www.kaggle.com/sreevishnudamodaran/build-custom-coco-annotations-512x512-tiled\n\nhttps://www.kaggle.com/sreevishnudamodaran/hubmap-coco-dataset-512x512-tiled\n\n\n#### Steps:\n\n**1. Organize all the files to a proper folder structure that is to be registered.**\nEx: \n```\n\t- images\n\t\t  - image_1\n\t\t  - image_2\n\t\t\t\t  .\n\t\t\t\t  .\n\t - masks\n\t\t   - mask_1\n\t\t   - mask_2\n\t\t\t\t  .\n\t\t\t\t  .\n```\nor\n\n```\n\t- coco_train\n\t\t- images\n\t\t\t- image_1\n\t\t\t- image_2\n\t\t\t\t   .\n\t\t\t\t   .\n\t\t- train.json (All annotations in coco format with proper relative path of images)\n```\n\t\t\n**2.  Create a Zip archive. Kaggle will automatically extract and unpack on uploading and registering the zip.**\n \n```python\n!zip -r dataset_name.zip ./coco_train\n```\n\n- Remove the unarchived folder if needed\n\n```python\n!rm -R coco_train\n```\n\n**3.  Check the new file size in MBs**\n\n```python\n!ls -ahl\n```\n\n**4. Click 'Save Version' to save the notebook and in the versions section, click 'Go to viewer' .**\n\n**5. Go to the bottom and click 'New Dataset'. This may be a little tricky to find so here's a screenshot.**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5532218%2F1567943cce91cff8fc4e305fc26c9c33%2Fstep.png?generation=1605948034551473&alt=media)\n\n\n\n\n\n\n**7. Add details, save and then go the Data section to see your dataset.**\n\n**8.Done. Add further details such as subtitle, metadata, tags, license etc as needed from the datasets page.**\n\n\n\n\n\n\n## UPDATE:\n## Alternate Method using Kaggle API directly from Notebook\n\n**1. Organize into data into proper folders as specified in the step 1 above**\n\n**2. Run the below cell with the right kaggle_username & kaggle_api_key. Get the kaggle_api_key from the 'Account' tab of your user profile (https://www.kaggle.com/<username>/account) and select 'Create API Token'. This will trigger the download of kaggle.json, a file containing your API credentials.**\n \n```\n%%bash\ntouch /root/.kaggle/kaggle.json\necho '{\"username\":\"<kaggle_username>\",\"key\":\"<kaggle_api_key>\"}' > /root/.kaggle/kaggle.json\ncat /root/.kaggle/kaggle.json\nchmod 600 /root/.kaggle/kaggle.json\n```\n\n**3. Test if the API works by listing datasets.**\n\n```\n!kaggle datasets list -s \"<search_word>\"\n```\n\n**4. Generate the dataset-metadata.json**\n\n```\n!kaggle datasets init -p <dataset_folder_name>\n```\n\n**5. Edit the metadata in the file.**\n\n```\nimport json\nf = open('./<dataset_folder_name>/dataset-metadata.json')\nmeta = json.load(f)\n\nmeta['title'] = \"<dataset_title>\"\nmeta['id'] = \"<kaggle_username>/<dataset_name>\"\n\nprint(meta)\n\nwith open('<dataset_folder_name>/dataset-metadata.json', 'w') as json_file:\n  json.dump(meta, json_file, indent=4)\n  \n```\n\n\n**6. Start the Dataset creation and upload.**\n\n```\n!kaggle datasets create -p ./<dataset_folder_name>/ -r zip\n```\n\nNote:\n- Specify 'zip' mode for directory upload with compression. Needed only if there are sub-directories within the dataset directory\n- It might take a few minutes to see an output.\n\n#### Hope everyone finds this useful."
  }
}