{
  "id": 271754,
  "title": "The data is too large and takes a very long time to process.",
  "url": "/competitions/landmark-recognition-2021/discussion/271754",
  "author_name": "",
  "post_date": "2021-09-12T12:47:13.917341Z",
  "votes": null,
  "comment_count": 11,
  "views": 0,
  "content": "<p>There are many large data competitions, but this competition requires a lot of computational resources. I'm processing GPU / TPU Limit every week, but I can't proceed at all. How are you guys doing?😄</p>",
  "messages": [
    {
      "id": "1510461",
      "postDate": "09/12/2021 12:47:13",
      "content": "<p>There are many large data competitions, but this competition requires a lot of computational resources. I'm processing GPU / TPU Limit every week, but I can't proceed at all. How are you guys doing?😄</p>",
      "rawMarkdown": "There are many large data competitions, but this competition requires a lot of computational resources. I'm processing GPU / TPU Limit every week, but I can't proceed at all. How are you guys doing?😄",
      "votes": null
    },
    {
      "id": "1510672",
      "postDate": "09/12/2021 16:32:45",
      "content": "<p>lol…same here… I am perpetually poor on TPU and GPU quota…😂..in fact my TPU quota runs over by the time its Monday morning…😆</p>",
      "rawMarkdown": "lol...same here... I am perpetually poor on TPU and GPU quota...😂..in fact my TPU quota runs over by the time its Monday morning...😆",
      "votes": null
    },
    {
      "id": "1510695",
      "postDate": "09/12/2021 16:47:46",
      "content": "<p>I suggest using Colab. Though it's using last gen TPU with lower memory, it's still faster compare to gpu. </p>",
      "rawMarkdown": "I suggest using Colab. Though it's using last gen TPU with lower memory, it's still faster compare to gpu.",
      "votes": null
    },
    {
      "id": "1511021",
      "postDate": "09/13/2021 03:22:08",
      "content": "<p>same as me. I alredy use TPU and GPU quota…. 😨</p>",
      "rawMarkdown": "same as me. I alredy use TPU and GPU quota.... 😨",
      "votes": null
    },
    {
      "id": "1511025",
      "postDate": "09/13/2021 03:24:15",
      "content": "<p>thank you your comment. I tried it. But Colab storage  is so small,so I alredy Give up.🤕</p>",
      "rawMarkdown": "thank you your comment. I tried it. But Colab storage  is so small,so I alredy Give up.🤕",
      "votes": null
    },
    {
      "id": "1511055",
      "postDate": "09/13/2021 04:39:02",
      "content": "<p>You won't need colab storage for anything other than storing models. The dataset can be stored in kaggle and can be accessed on colab TPU with the help of GCS path.</p>",
      "rawMarkdown": "You won't need colab storage for anything other than storing models. The dataset can be stored in kaggle and can be accessed on colab TPU with the help of GCS path.",
      "votes": null
    },
    {
      "id": "1511088",
      "postDate": "09/13/2021 05:24:33",
      "content": "<p>for example</p>\n<pre><code>data = KaggleDatasets().get_gcs_path(f\"landmark-recognition-2021-tfrecords-fold0\")\nprint(data)\n</code></pre>\n<p>run the code above in kaggle script, you'll get<br>\n<code>gs://kds-3a731c7599bde541904c88c4a7751fe46facd781aa156de959fc819a</code></p>\n<p>then you can simply copy that address in colad script as an path and read whatever is inside.<br>\none thing to notice though, the address might change from time to time.</p>",
      "rawMarkdown": "for example\n```python\ndata = KaggleDatasets().get_gcs_path(f\"landmark-recognition-2021-tfrecords-fold0\")\nprint(data)\n```\nrun the code above in kaggle script, you'll get\n`gs://kds-3a731c7599bde541904c88c4a7751fe46facd781aa156de959fc819a`\n\nthen you can simply copy that address in colad script as an path and read whatever is inside.\none thing to notice though, the address might change from time to time.",
      "votes": null
    },
    {
      "id": "1511248",
      "postDate": "09/13/2021 09:01:57",
      "content": "<p>Thanks for sharing this. Didn't know we can access Kaggle datasets directly from Colab without copying the tfrecords to Colab. Thanks for this tip. Much appreciated.</p>",
      "rawMarkdown": "Thanks for sharing this. Didn't know we can access Kaggle datasets directly from Colab without copying the tfrecords to Colab. Thanks for this tip. Much appreciated.",
      "votes": null
    },
    {
      "id": "1511317",
      "postDate": "09/13/2021 10:22:44",
      "content": "<p>Thank you sharering  I try to do that.😄</p>",
      "rawMarkdown": "Thank you sharering  I try to do that.😄",
      "votes": null
    },
    {
      "id": "1511360",
      "postDate": "09/13/2021 11:04:53",
      "content": "<p>I tried but error occured,<br>\nDoes anyone help me?</p>\n<hr>\n<p>from google.colab import auth<br>\nauth.authenticate_user()</p>\n<p>!echo \"deb <a href=\"http://packages.cloud.google.com/apt\" target=\"_blank\">http://packages.cloud.google.com/apt</a> gcsfuse-bionic main\" &gt; /etc/apt/sources.list.d/gcsfuse.list<br>\n!curl <a href=\"https://packages.cloud.google.com/apt/doc/apt-key.gpg\" target=\"_blank\">https://packages.cloud.google.com/apt/doc/apt-key.gpg</a> | apt-key add -<br>\n!apt update<br>\n!apt install gcsfuse</p>\n<p>!mkdir fold0<br>\nBACKET_NAME='gs://kds-00ff638eb371cb41f8a49ca90a8cd084e121e4eba48f691e7ef4507f'</p>\n<p>! gcsfuse --implicit-dirs --limit-bytes-per-sec -1 --limit-ops-per-sec -1 ${BUCKET_NAME} /content/fold0<br>\n2021/09/13 11:01:35.496502 Using mount point: /content/fold0<br>\n2021/09/13 11:01:35.506358 Opening GCS connection…<br>\n2021/09/13 11:01:35.507408 Mounting file system \"gcsfuse\"…<br>\n2021/09/13 11:01:35.509961 File system has been successfully mounted.</p>\n<hr>\n<p>!ls ./fold0<br>\nls: reading directory './fold0': Input/output error</p>",
      "rawMarkdown": "I tried but error occured,\nDoes anyone help me?\n\n----------------------\nfrom google.colab import auth\nauth.authenticate_user()\n\n!echo \"deb http://packages.cloud.google.com/apt gcsfuse-bionic main\" > /etc/apt/sources.list.d/gcsfuse.list\n!curl https://packages.cloud.google.com/apt/doc/apt-key.gpg | apt-key add -\n!apt update\n!apt install gcsfuse\n\n!mkdir fold0\nBACKET_NAME='gs://kds-00ff638eb371cb41f8a49ca90a8cd084e121e4eba48f691e7ef4507f'\n\n! gcsfuse --implicit-dirs --limit-bytes-per-sec -1 --limit-ops-per-sec -1 ${BUCKET_NAME} /content/fold0\n2021/09/13 11:01:35.496502 Using mount point: /content/fold0\n2021/09/13 11:01:35.506358 Opening GCS connection...\n2021/09/13 11:01:35.507408 Mounting file system \"gcsfuse\"...\n2021/09/13 11:01:35.509961 File system has been successfully mounted.\n\n\n------------------------\n\n!ls ./fold0\nls: reading directory './fold0': Input/output error",
      "votes": null
    },
    {
      "id": "1512610",
      "postDate": "09/14/2021 12:09:46",
      "content": "<p><a href=\"https://www.kaggle.com/tensorchoko\" target=\"_blank\">@tensorchoko</a> <br>\ndon't think <code>ls</code> will work here.<br>\nWhat I meant is more something like this.</p>\n<pre><code>path = 'gs://kds-00ff638eb371cb41f8a49ca90a8cd084e121e4eba48f691e7ef4507f'\ndataset = tf.data.TFRecordDataset(tf.io.gfile.glob(path + \"/*.tfrec\"), num_parallel_reads=AUTO)\n</code></pre>",
      "rawMarkdown": "tensorchoko \ndon't think `ls` will work here.\nWhat I meant is more something like this.\n```python\npath = 'gs://kds-00ff638eb371cb41f8a49ca90a8cd084e121e4eba48f691e7ef4507f'\ndataset = tf.data.TFRecordDataset(tf.io.gfile.glob(path + \"/*.tfrec\"), num_parallel_reads=AUTO)\n```",
      "votes": null
    },
    {
      "id": "1513171",
      "postDate": "09/14/2021 22:49:31",
      "content": "<p>Thanks. I'll try again</p>",
      "rawMarkdown": "Thanks. I'll try again",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1510672,
      "author_name": "sandy1112",
      "author_url": "",
      "post_date": "09/12/2021 16:32:45",
      "content": "<p>lol…same here… I am perpetually poor on TPU and GPU quota…😂..in fact my TPU quota runs over by the time its Monday morning…😆</p>",
      "votes": null,
      "replies": [
        {
          "id": 1511021,
          "author_name": "tensorchoko",
          "author_url": "",
          "post_date": "09/13/2021 03:22:08",
          "content": "<p>same as me. I alredy use TPU and GPU quota…. 😨</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1510695,
      "author_name": "gdoong",
      "author_url": "",
      "post_date": "09/12/2021 16:47:46",
      "content": "<p>I suggest using Colab. Though it's using last gen TPU with lower memory, it's still faster compare to gpu. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1511025,
          "author_name": "tensorchoko",
          "author_url": "",
          "post_date": "09/13/2021 03:24:15",
          "content": "<p>thank you your comment. I tried it. But Colab storage  is so small,so I alredy Give up.🤕</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1511055,
          "author_name": "ks2019",
          "author_url": "",
          "post_date": "09/13/2021 04:39:02",
          "content": "<p>You won't need colab storage for anything other than storing models. The dataset can be stored in kaggle and can be accessed on colab TPU with the help of GCS path.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1511088,
          "author_name": "gdoong",
          "author_url": "",
          "post_date": "09/13/2021 05:24:33",
          "content": "<p>for example</p>\n<pre><code>data = KaggleDatasets().get_gcs_path(f\"landmark-recognition-2021-tfrecords-fold0\")\nprint(data)\n</code></pre>\n<p>run the code above in kaggle script, you'll get<br>\n<code>gs://kds-3a731c7599bde541904c88c4a7751fe46facd781aa156de959fc819a</code></p>\n<p>then you can simply copy that address in colad script as an path and read whatever is inside.<br>\none thing to notice though, the address might change from time to time.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1511248,
          "author_name": "sandy1112",
          "author_url": "",
          "post_date": "09/13/2021 09:01:57",
          "content": "<p>Thanks for sharing this. Didn't know we can access Kaggle datasets directly from Colab without copying the tfrecords to Colab. Thanks for this tip. Much appreciated.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1511317,
          "author_name": "tensorchoko",
          "author_url": "",
          "post_date": "09/13/2021 10:22:44",
          "content": "<p>Thank you sharering  I try to do that.😄</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1511360,
          "author_name": "tensorchoko",
          "author_url": "",
          "post_date": "09/13/2021 11:04:53",
          "content": "<p>I tried but error occured,<br>\nDoes anyone help me?</p>\n<hr>\n<p>from google.colab import auth<br>\nauth.authenticate_user()</p>\n<p>!echo \"deb <a href=\"http://packages.cloud.google.com/apt\" target=\"_blank\">http://packages.cloud.google.com/apt</a> gcsfuse-bionic main\" &gt; /etc/apt/sources.list.d/gcsfuse.list<br>\n!curl <a href=\"https://packages.cloud.google.com/apt/doc/apt-key.gpg\" target=\"_blank\">https://packages.cloud.google.com/apt/doc/apt-key.gpg</a> | apt-key add -<br>\n!apt update<br>\n!apt install gcsfuse</p>\n<p>!mkdir fold0<br>\nBACKET_NAME='gs://kds-00ff638eb371cb41f8a49ca90a8cd084e121e4eba48f691e7ef4507f'</p>\n<p>! gcsfuse --implicit-dirs --limit-bytes-per-sec -1 --limit-ops-per-sec -1 ${BUCKET_NAME} /content/fold0<br>\n2021/09/13 11:01:35.496502 Using mount point: /content/fold0<br>\n2021/09/13 11:01:35.506358 Opening GCS connection…<br>\n2021/09/13 11:01:35.507408 Mounting file system \"gcsfuse\"…<br>\n2021/09/13 11:01:35.509961 File system has been successfully mounted.</p>\n<hr>\n<p>!ls ./fold0<br>\nls: reading directory './fold0': Input/output error</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1512610,
          "author_name": "gdoong",
          "author_url": "",
          "post_date": "09/14/2021 12:09:46",
          "content": "<p><a href=\"https://www.kaggle.com/tensorchoko\" target=\"_blank\">@tensorchoko</a> <br>\ndon't think <code>ls</code> will work here.<br>\nWhat I meant is more something like this.</p>\n<pre><code>path = 'gs://kds-00ff638eb371cb41f8a49ca90a8cd084e121e4eba48f691e7ef4507f'\ndataset = tf.data.TFRecordDataset(tf.io.gfile.glob(path + \"/*.tfrec\"), num_parallel_reads=AUTO)\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1513171,
          "author_name": "tensorchoko",
          "author_url": "",
          "post_date": "09/14/2021 22:49:31",
          "content": "<p>Thanks. I'll try again</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1510461": "There are many large data competitions, but this competition requires a lot of computational resources. I'm processing GPU / TPU Limit every week, but I can't proceed at all. How are you guys doing?😄",
    "1510672": "lol...same here... I am perpetually poor on TPU and GPU quota...😂..in fact my TPU quota runs over by the time its Monday morning...😆",
    "1510695": "I suggest using Colab. Though it's using last gen TPU with lower memory, it's still faster compare to gpu.",
    "1511021": "same as me. I alredy use TPU and GPU quota.... 😨",
    "1511025": "thank you your comment. I tried it. But Colab storage  is so small,so I alredy Give up.🤕",
    "1511055": "You won't need colab storage for anything other than storing models. The dataset can be stored in kaggle and can be accessed on colab TPU with the help of GCS path.",
    "1511088": "for example\n```python\ndata = KaggleDatasets().get_gcs_path(f\"landmark-recognition-2021-tfrecords-fold0\")\nprint(data)\n```\nrun the code above in kaggle script, you'll get\n`gs://kds-3a731c7599bde541904c88c4a7751fe46facd781aa156de959fc819a`\n\nthen you can simply copy that address in colad script as an path and read whatever is inside.\none thing to notice though, the address might change from time to time.",
    "1511248": "Thanks for sharing this. Didn't know we can access Kaggle datasets directly from Colab without copying the tfrecords to Colab. Thanks for this tip. Much appreciated.",
    "1511317": "Thank you sharering  I try to do that.😄",
    "1511360": "I tried but error occured,\nDoes anyone help me?\n\n----------------------\nfrom google.colab import auth\nauth.authenticate_user()\n\n!echo \"deb http://packages.cloud.google.com/apt gcsfuse-bionic main\" > /etc/apt/sources.list.d/gcsfuse.list\n!curl https://packages.cloud.google.com/apt/doc/apt-key.gpg | apt-key add -\n!apt update\n!apt install gcsfuse\n\n!mkdir fold0\nBACKET_NAME='gs://kds-00ff638eb371cb41f8a49ca90a8cd084e121e4eba48f691e7ef4507f'\n\n! gcsfuse --implicit-dirs --limit-bytes-per-sec -1 --limit-ops-per-sec -1 ${BUCKET_NAME} /content/fold0\n2021/09/13 11:01:35.496502 Using mount point: /content/fold0\n2021/09/13 11:01:35.506358 Opening GCS connection...\n2021/09/13 11:01:35.507408 Mounting file system \"gcsfuse\"...\n2021/09/13 11:01:35.509961 File system has been successfully mounted.\n\n\n------------------------\n\n!ls ./fold0\nls: reading directory './fold0': Input/output error",
    "1512610": "tensorchoko \ndon't think `ls` will work here.\nWhat I meant is more something like this.\n```python\npath = 'gs://kds-00ff638eb371cb41f8a49ca90a8cd084e121e4eba48f691e7ef4507f'\ndataset = tf.data.TFRecordDataset(tf.io.gfile.glob(path + \"/*.tfrec\"), num_parallel_reads=AUTO)\n```",
    "1513171": "Thanks. I'll try again"
  },
  "source": "meta"
}