{
  "id": 231373,
  "title": "[HELP] Unable to fetch GCS Path of Kaggle Datasets",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/231373",
  "author_name": "Satwik",
  "post_date": "2021-04-08T08:01:12.886000",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I have been trying to use TPU to train , and getting GCS Paths of datasets isn't working only for a few SPECIFIC Datasets. It's working fine for the rest. For example , </p>\n<p><code>BLUE_TILE_DIR = KaggleDatasets().get_gcs_path(\"human-protein-atlas-blue-cell-tile-dataset\")</code></p>\n<p>works perfectly fine. But in the same notebook , </p>\n<p><code>YELLOW_TILE_DIR = KaggleDatasets().get_gcs_path('human-protein-atlas-yellow-cell-tile-dataset')</code></p>\n<p>gives the following error -</p>\n<pre><code>---------------------------------------------------------------------------\nHTTPError                                 Traceback (most recent call last)\n~/.local/lib/python3.7/site-packages/kaggle_web_client.py in make_post_request(self, data, endpoint, timeout)\n     45         try:\n---&gt; 46             with urllib.request.urlopen(req, timeout=timeout) as response:\n     47                 response_json = json.loads(response.read())\n\n/opt/conda/lib/python3.7/urllib/request.py in urlopen(url, data, timeout, cafile, capath, cadefault, context)\n    221         opener = _opener\n--&gt; 222     return opener.open(url, data, timeout)\n    223 \n\n/opt/conda/lib/python3.7/urllib/request.py in open(self, fullurl, data, timeout)\n    530             meth = getattr(processor, meth_name)\n--&gt; 531             response = meth(req, response)\n    532 \n\n/opt/conda/lib/python3.7/urllib/request.py in http_response(self, request, response)\n    640             response = self.parent.error(\n--&gt; 641                 'http', request, response, code, msg, hdrs)\n    642 \n\n/opt/conda/lib/python3.7/urllib/request.py in error(self, proto, *args)\n    568             args = (dict, 'default', 'http_error_default') + orig_args\n--&gt; 569             return self._call_chain(*args)\n    570 \n\n/opt/conda/lib/python3.7/urllib/request.py in _call_chain(self, chain, kind, meth_name, *args)\n    502             func = getattr(handler, meth_name)\n--&gt; 503             result = func(*args)\n    504             if result is not None:\n\n/opt/conda/lib/python3.7/urllib/request.py in http_error_default(self, req, fp, code, msg, hdrs)\n    648     def http_error_default(self, req, fp, code, msg, hdrs):\n--&gt; 649         raise HTTPError(req.full_url, code, msg, hdrs, fp)\n    650 \n\nHTTPError: HTTP Error 502: Bad Gateway\n\nThe above exception was the direct cause of the following exception:\n\nConnectionError                           Traceback (most recent call last)\n&lt;ipython-input-11-85cfb3a1672e&gt; in &lt;module&gt;\n----&gt; 1 YELLOW_TILE_DIR = KaggleDatasets().get_gcs_path('human-protein-atlas-yellow-cell-tile-dataset')\n\n~/.local/lib/python3.7/site-packages/kaggle_datasets.py in get_gcs_path(self, dataset_dir)\n     22             'IntegrationType': integration_type,\n     23         }\n---&gt; 24         result = self.web_client.make_post_request(data, self.GET_GCS_PATH_ENDPOINT, self.TIMEOUT_SECS)\n     25         return result['destinationBucket']\n\n~/.local/lib/python3.7/site-packages/kaggle_web_client.py in make_post_request(self, data, endpoint, timeout)\n     57                     'Timeout error trying to communicate with service. Please ensure internet is on.') from e\n     58             raise ConnectionError(\n---&gt; 59                 'Connection error trying to communicate with service.') from e\n     60         except HTTPError as e:\n     61             if e.code == 401 or e.code == 403:\n\nConnectionError: Connection error trying to communicate with service.\n</code></pre>\n<p>Internet is on ,  TPU is enabled. Datasets are added to the notebook's input. This is really a huge block for me , please let me know if there are any workarounds to this. Thank you!</p>",
  "messages": [
    {
      "id": 1267071,
      "postDate": "2021-04-08T09:50:58.397Z",
      "content": "<p>Update - Restarting internet on the notebook managed to make it work. However now , I am not able to open any files in the GS Dataset. I double checked the path , it is correct. Tried using gsutil ls on the GS path , it is listing all the files present in the directory properly. But when I try to open a single file from the path , it says file not found. :(</p>\n<p>Update 2 - Fixed the issue. I was reading the files from the GCS bucket, using PIL to open the files. Which is not really a good way to read files from GCS ,  once I used tf.io.read_file(),  it was working properly.</p>",
      "rawMarkdown": "Update - Restarting internet on the notebook managed to make it work. However now , I am not able to open any files in the GS Dataset. I double checked the path , it is correct. Tried using gsutil ls on the GS path , it is listing all the files present in the directory properly. But when I try to open a single file from the path , it says file not found. :(\n\nUpdate 2 - Fixed the issue. I was reading the files from the GCS bucket, using PIL to open the files. Which is not really a good way to read files from GCS ,  once I used tf.io.read_file(),  it was working properly.",
      "votes": 2
    },
    {
      "id": 1266969,
      "postDate": "2021-04-08T08:01:12.887Z",
      "content": "<p>I have been trying to use TPU to train , and getting GCS Paths of datasets isn't working only for a few SPECIFIC Datasets. It's working fine for the rest. For example , </p>\n<p><code>BLUE_TILE_DIR = KaggleDatasets().get_gcs_path(\"human-protein-atlas-blue-cell-tile-dataset\")</code></p>\n<p>works perfectly fine. But in the same notebook , </p>\n<p><code>YELLOW_TILE_DIR = KaggleDatasets().get_gcs_path('human-protein-atlas-yellow-cell-tile-dataset')</code></p>\n<p>gives the following error -</p>\n<pre><code>---------------------------------------------------------------------------\nHTTPError                                 Traceback (most recent call last)\n~/.local/lib/python3.7/site-packages/kaggle_web_client.py in make_post_request(self, data, endpoint, timeout)\n     45         try:\n---&gt; 46             with urllib.request.urlopen(req, timeout=timeout) as response:\n     47                 response_json = json.loads(response.read())\n\n/opt/conda/lib/python3.7/urllib/request.py in urlopen(url, data, timeout, cafile, capath, cadefault, context)\n    221         opener = _opener\n--&gt; 222     return opener.open(url, data, timeout)\n    223 \n\n/opt/conda/lib/python3.7/urllib/request.py in open(self, fullurl, data, timeout)\n    530             meth = getattr(processor, meth_name)\n--&gt; 531             response = meth(req, response)\n    532 \n\n/opt/conda/lib/python3.7/urllib/request.py in http_response(self, request, response)\n    640             response = self.parent.error(\n--&gt; 641                 'http', request, response, code, msg, hdrs)\n    642 \n\n/opt/conda/lib/python3.7/urllib/request.py in error(self, proto, *args)\n    568             args = (dict, 'default', 'http_error_default') + orig_args\n--&gt; 569             return self._call_chain(*args)\n    570 \n\n/opt/conda/lib/python3.7/urllib/request.py in _call_chain(self, chain, kind, meth_name, *args)\n    502             func = getattr(handler, meth_name)\n--&gt; 503             result = func(*args)\n    504             if result is not None:\n\n/opt/conda/lib/python3.7/urllib/request.py in http_error_default(self, req, fp, code, msg, hdrs)\n    648     def http_error_default(self, req, fp, code, msg, hdrs):\n--&gt; 649         raise HTTPError(req.full_url, code, msg, hdrs, fp)\n    650 \n\nHTTPError: HTTP Error 502: Bad Gateway\n\nThe above exception was the direct cause of the following exception:\n\nConnectionError                           Traceback (most recent call last)\n&lt;ipython-input-11-85cfb3a1672e&gt; in &lt;module&gt;\n----&gt; 1 YELLOW_TILE_DIR = KaggleDatasets().get_gcs_path('human-protein-atlas-yellow-cell-tile-dataset')\n\n~/.local/lib/python3.7/site-packages/kaggle_datasets.py in get_gcs_path(self, dataset_dir)\n     22             'IntegrationType': integration_type,\n     23         }\n---&gt; 24         result = self.web_client.make_post_request(data, self.GET_GCS_PATH_ENDPOINT, self.TIMEOUT_SECS)\n     25         return result['destinationBucket']\n\n~/.local/lib/python3.7/site-packages/kaggle_web_client.py in make_post_request(self, data, endpoint, timeout)\n     57                     'Timeout error trying to communicate with service. Please ensure internet is on.') from e\n     58             raise ConnectionError(\n---&gt; 59                 'Connection error trying to communicate with service.') from e\n     60         except HTTPError as e:\n     61             if e.code == 401 or e.code == 403:\n\nConnectionError: Connection error trying to communicate with service.\n</code></pre>\n<p>Internet is on ,  TPU is enabled. Datasets are added to the notebook's input. This is really a huge block for me , please let me know if there are any workarounds to this. Thank you!</p>",
      "rawMarkdown": "I have been trying to use TPU to train , and getting GCS Paths of datasets isn't working only for a few SPECIFIC Datasets. It's working fine for the rest. For example , \n\n`BLUE_TILE_DIR = KaggleDatasets().get_gcs_path(\"human-protein-atlas-blue-cell-tile-dataset\")`\n\nworks perfectly fine. But in the same notebook , \n\n`YELLOW_TILE_DIR = KaggleDatasets().get_gcs_path('human-protein-atlas-yellow-cell-tile-dataset')`\n\ngives the following error -\n\n```\n---------------------------------------------------------------------------\nHTTPError                                 Traceback (most recent call last)\n~/.local/lib/python3.7/site-packages/kaggle_web_client.py in make_post_request(self, data, endpoint, timeout)\n     45         try:\n---> 46             with urllib.request.urlopen(req, timeout=timeout) as response:\n     47                 response_json = json.loads(response.read())\n\n/opt/conda/lib/python3.7/urllib/request.py in urlopen(url, data, timeout, cafile, capath, cadefault, context)\n    221         opener = _opener\n--> 222     return opener.open(url, data, timeout)\n    223 \n\n/opt/conda/lib/python3.7/urllib/request.py in open(self, fullurl, data, timeout)\n    530             meth = getattr(processor, meth_name)\n--> 531             response = meth(req, response)\n    532 \n\n/opt/conda/lib/python3.7/urllib/request.py in http_response(self, request, response)\n    640             response = self.parent.error(\n--> 641                 'http', request, response, code, msg, hdrs)\n    642 \n\n/opt/conda/lib/python3.7/urllib/request.py in error(self, proto, *args)\n    568             args = (dict, 'default', 'http_error_default') + orig_args\n--> 569             return self._call_chain(*args)\n    570 \n\n/opt/conda/lib/python3.7/urllib/request.py in _call_chain(self, chain, kind, meth_name, *args)\n    502             func = getattr(handler, meth_name)\n--> 503             result = func(*args)\n    504             if result is not None:\n\n/opt/conda/lib/python3.7/urllib/request.py in http_error_default(self, req, fp, code, msg, hdrs)\n    648     def http_error_default(self, req, fp, code, msg, hdrs):\n--> 649         raise HTTPError(req.full_url, code, msg, hdrs, fp)\n    650 \n\nHTTPError: HTTP Error 502: Bad Gateway\n\nThe above exception was the direct cause of the following exception:\n\nConnectionError                           Traceback (most recent call last)\n<ipython-input-11-85cfb3a1672e> in <module>\n----> 1 YELLOW_TILE_DIR = KaggleDatasets().get_gcs_path('human-protein-atlas-yellow-cell-tile-dataset')\n\n~/.local/lib/python3.7/site-packages/kaggle_datasets.py in get_gcs_path(self, dataset_dir)\n     22             'IntegrationType': integration_type,\n     23         }\n---> 24         result = self.web_client.make_post_request(data, self.GET_GCS_PATH_ENDPOINT, self.TIMEOUT_SECS)\n     25         return result['destinationBucket']\n\n~/.local/lib/python3.7/site-packages/kaggle_web_client.py in make_post_request(self, data, endpoint, timeout)\n     57                     'Timeout error trying to communicate with service. Please ensure internet is on.') from e\n     58             raise ConnectionError(\n---> 59                 'Connection error trying to communicate with service.') from e\n     60         except HTTPError as e:\n     61             if e.code == 401 or e.code == 403:\n\nConnectionError: Connection error trying to communicate with service.\n```\n\n\nInternet is on ,  TPU is enabled. Datasets are added to the notebook's input. This is really a huge block for me , please let me know if there are any workarounds to this. Thank you!",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 1267071,
      "author_name": "Satwik",
      "author_url": "",
      "post_date": "2021-04-08T09:50:58.397000",
      "content": "<p>Update - Restarting internet on the notebook managed to make it work. However now , I am not able to open any files in the GS Dataset. I double checked the path , it is correct. Tried using gsutil ls on the GS path , it is listing all the files present in the directory properly. But when I try to open a single file from the path , it says file not found. :(</p>\n<p>Update 2 - Fixed the issue. I was reading the files from the GCS bucket, using PIL to open the files. Which is not really a good way to read files from GCS ,  once I used tf.io.read_file(),  it was working properly.</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1267071": "Update - Restarting internet on the notebook managed to make it work. However now , I am not able to open any files in the GS Dataset. I double checked the path , it is correct. Tried using gsutil ls on the GS path , it is listing all the files present in the directory properly. But when I try to open a single file from the path , it says file not found. :(\n\nUpdate 2 - Fixed the issue. I was reading the files from the GCS bucket, using PIL to open the files. Which is not really a good way to read files from GCS ,  once I used tf.io.read_file(),  it was working properly.",
    "1266969": "I have been trying to use TPU to train , and getting GCS Paths of datasets isn't working only for a few SPECIFIC Datasets. It's working fine for the rest. For example , \n\n`BLUE_TILE_DIR = KaggleDatasets().get_gcs_path(\"human-protein-atlas-blue-cell-tile-dataset\")`\n\nworks perfectly fine. But in the same notebook , \n\n`YELLOW_TILE_DIR = KaggleDatasets().get_gcs_path('human-protein-atlas-yellow-cell-tile-dataset')`\n\ngives the following error -\n\n```\n---------------------------------------------------------------------------\nHTTPError                                 Traceback (most recent call last)\n~/.local/lib/python3.7/site-packages/kaggle_web_client.py in make_post_request(self, data, endpoint, timeout)\n     45         try:\n---> 46             with urllib.request.urlopen(req, timeout=timeout) as response:\n     47                 response_json = json.loads(response.read())\n\n/opt/conda/lib/python3.7/urllib/request.py in urlopen(url, data, timeout, cafile, capath, cadefault, context)\n    221         opener = _opener\n--> 222     return opener.open(url, data, timeout)\n    223 \n\n/opt/conda/lib/python3.7/urllib/request.py in open(self, fullurl, data, timeout)\n    530             meth = getattr(processor, meth_name)\n--> 531             response = meth(req, response)\n    532 \n\n/opt/conda/lib/python3.7/urllib/request.py in http_response(self, request, response)\n    640             response = self.parent.error(\n--> 641                 'http', request, response, code, msg, hdrs)\n    642 \n\n/opt/conda/lib/python3.7/urllib/request.py in error(self, proto, *args)\n    568             args = (dict, 'default', 'http_error_default') + orig_args\n--> 569             return self._call_chain(*args)\n    570 \n\n/opt/conda/lib/python3.7/urllib/request.py in _call_chain(self, chain, kind, meth_name, *args)\n    502             func = getattr(handler, meth_name)\n--> 503             result = func(*args)\n    504             if result is not None:\n\n/opt/conda/lib/python3.7/urllib/request.py in http_error_default(self, req, fp, code, msg, hdrs)\n    648     def http_error_default(self, req, fp, code, msg, hdrs):\n--> 649         raise HTTPError(req.full_url, code, msg, hdrs, fp)\n    650 \n\nHTTPError: HTTP Error 502: Bad Gateway\n\nThe above exception was the direct cause of the following exception:\n\nConnectionError                           Traceback (most recent call last)\n<ipython-input-11-85cfb3a1672e> in <module>\n----> 1 YELLOW_TILE_DIR = KaggleDatasets().get_gcs_path('human-protein-atlas-yellow-cell-tile-dataset')\n\n~/.local/lib/python3.7/site-packages/kaggle_datasets.py in get_gcs_path(self, dataset_dir)\n     22             'IntegrationType': integration_type,\n     23         }\n---> 24         result = self.web_client.make_post_request(data, self.GET_GCS_PATH_ENDPOINT, self.TIMEOUT_SECS)\n     25         return result['destinationBucket']\n\n~/.local/lib/python3.7/site-packages/kaggle_web_client.py in make_post_request(self, data, endpoint, timeout)\n     57                     'Timeout error trying to communicate with service. Please ensure internet is on.') from e\n     58             raise ConnectionError(\n---> 59                 'Connection error trying to communicate with service.') from e\n     60         except HTTPError as e:\n     61             if e.code == 401 or e.code == 403:\n\nConnectionError: Connection error trying to communicate with service.\n```\n\n\nInternet is on ,  TPU is enabled. Datasets are added to the notebook's input. This is really a huge block for me , please let me know if there are any workarounds to this. Thank you!"
  }
}