{
  "id": 143770,
  "title": "When will we need GCS_DS_PATH = KaggleDatasets().get_gcs_path ?",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/143770",
  "author_name": "",
  "post_date": "2020-04-16T07:53:15.641890400Z",
  "votes": 6,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi all, I am a newbie on TPU.</p>\n\n<p>I am studying <a href=\"/xhlulu\">@xhlulu</a> great kernel and also other TPU kernels in image competitions.\n<a href=\"https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\">https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta</a></p>\n\n<p>In competitions like Flower or Plant, it seems we need to use this command:\n<code>GCS_DS_PATH = KaggleDatasets().get_gcs_path()</code></p>\n\n<p>I am a bit confused why we don't need it here.</p>\n\n<p>Thanks!!</p>",
  "messages": [
    {
      "id": "809448",
      "postDate": "04/16/2020 07:53:15",
      "content": "<p>Hi all, I am a newbie on TPU.</p>\n\n<p>I am studying <a href=\"/xhlulu\">@xhlulu</a> great kernel and also other TPU kernels in image competitions.\n<a href=\"https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\">https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta</a></p>\n\n<p>In competitions like Flower or Plant, it seems we need to use this command:\n<code>GCS_DS_PATH = KaggleDatasets().get_gcs_path()</code></p>\n\n<p>I am a bit confused why we don't need it here.</p>\n\n<p>Thanks!!</p>",
      "rawMarkdown": "Hi all, I am a newbie on TPU.\n\nI am studying @xhlulu great kernel and also other TPU kernels in image competitions.\nhttps://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\n\nIn competitions like Flower or Plant, it seems we need to use this command:\n`GCS_DS_PATH = KaggleDatasets().get_gcs_path()`\n\nI am a bit confused why we don't need it here.\n\nThanks!!",
      "votes": null
    },
    {
      "id": "809455",
      "postDate": "04/16/2020 07:56:59",
      "content": "<p>Because we are using data-set from local drives?</p>",
      "rawMarkdown": "Because we are using data-set from local drives?",
      "votes": null
    },
    {
      "id": "809471",
      "postDate": "04/16/2020 08:09:21",
      "content": "<p>I am not sure, Aditya. </p>\n\n<p>In fact, it's a bit off-topic from this comp., but just for reference,\nI am playing around with TF2.1 image captioning for fun .. I copy all the codes from TF2.1 website, and luckily use Flickr / COCO dataset in locally located in Kaggle (instead of downloading using internet like in the example)</p>\n\n<p><a href=\"https://www.kaggle.com/ratthachat/image-captioning-using-attention/\">https://www.kaggle.com/ratthachat/image-captioning-using-attention/</a> (made public just for reference -- this kernel is still very messy)</p>\n\n<p>But in this case, it seems I have to use GCS_DS_PATH = KaggleDatasets().get_gcs_path() \n(even thought the data is in local drive too)</p>\n\n<p>And one more question, if somebody know :D ,  is there a limit on dataset size to use <code>get_gcs_path()</code> (e.g. I use full COCO2014 with 18GB and it got HTTP error)</p>",
      "rawMarkdown": "I am not sure, Aditya. \n\nIn fact, it's a bit off-topic from this comp., but just for reference,\nI am playing around with TF2.1 image captioning for fun .. I copy all the codes from TF2.1 website, and luckily use Flickr / COCO dataset in locally located in Kaggle (instead of downloading using internet like in the example)\n\nhttps://www.kaggle.com/ratthachat/image-captioning-using-attention/ (made public just for reference -- this kernel is still very messy)\n\nBut in this case, it seems I have to use GCS_DS_PATH = KaggleDatasets().get_gcs_path() \n(even thought the data is in local drive too)\n\nAnd one more question, if somebody know :D ,  is there a limit on dataset size to use `get_gcs_path()` (e.g. I use full COCO2014 with 18GB and it got HTTP error)",
      "votes": null
    },
    {
      "id": "809479",
      "postDate": "04/16/2020 08:15:27",
      "content": "<p>Hmm interesting, IIRC, they move the data to GCS nearest to the VM we use automatically? (i don't remember where i read something similar to this;)</p>",
      "rawMarkdown": "Hmm interesting, IIRC, they move the data to GCS nearest to the VM we use automatically? (i don't remember where i read something similar to this;)",
      "votes": null
    },
    {
      "id": "810149",
      "postDate": "04/16/2020 18:50:27",
      "content": "<p>If you training on data in memory, you can load the data any way you want.</p>\n\n<p>If you are using the tf.data.Dataset API to work with a dataset that does not fit in memory, then your data pipeline will be shipped to the TPU and executed there. The TPU will be doing the loading and the TPU can access data in GCS (Google Cloud Storage) only.</p>\n\n<p>Fortunately, Kaggle can get you a pointer to the competition data in GCS through:\n<code>\nGCS_DS_PATH = KaggleDatasets().get_gcs_path()\n</code>\nAs an added bonus, Kaggle will give you a GCS bucket that is always colocated with the TPU so you don't end up running in Europe and loading data from the west coast. That would be slow.</p>",
      "rawMarkdown": "If you training on data in memory, you can load the data any way you want.\n\nIf you are using the tf.data.Dataset API to work with a dataset that does not fit in memory, then your data pipeline will be shipped to the TPU and executed there. The TPU will be doing the loading and the TPU can access data in GCS (Google Cloud Storage) only.\n\nFortunately, Kaggle can get you a pointer to the competition data in GCS through:\n```\nGCS_DS_PATH = KaggleDatasets().get_gcs_path()\n```\nAs an added bonus, Kaggle will give you a GCS bucket that is always colocated with the TPU so you don't end up running in Europe and loading data from the west coast. That would be slow.",
      "votes": null
    },
    {
      "id": "810151",
      "postDate": "04/16/2020 18:53:11",
      "content": "<p>For data in memory, TPU training on numpy arrays directly works (in Keras model.fit()). Also, any memory array can be converted to a tf.data.Dataset with <code>Dataset.from_tensor_slices()</code> and then you can train on the tf.data.Dataset directly.</p>",
      "rawMarkdown": "For data in memory, TPU training on numpy arrays directly works (in Keras model.fit()). Also, any memory array can be converted to a tf.data.Dataset with `Dataset.from_tensor_slices()` and then you can train on the tf.data.Dataset directly.",
      "votes": null
    },
    {
      "id": "810424",
      "postDate": "04/17/2020 00:05:50",
      "content": "<p>Thank you for explanation, Martin! <a href=\"/mgornergoogle\">@mgornergoogle</a> \nIs there any limitation on the size of the dataset to be used with <code>KaggleDatasets().get_gcs_path()</code> ?</p>\n\n<p>E.g. I tried a dataset of size 10-19GB and got an HTTP error. (internet is on)</p>",
      "rawMarkdown": "Thank you for explanation, Martin! @mgornergoogle \nIs there any limitation on the size of the dataset to be used with `KaggleDatasets().get_gcs_path()` ?\n\nE.g. I tried a dataset of size 10-19GB and got an HTTP error. (internet is on)",
      "votes": null
    },
    {
      "id": "810508",
      "postDate": "04/17/2020 03:15:01",
      "content": "<p>Try again. Yes, in the current implementation it can time out on a large dataset the first time someone requests it, but the  bucket creation and data transfer still complete in the background. The next time you try, it will just return the name of the existing bucket.</p>",
      "rawMarkdown": "Try again. Yes, in the current implementation it can time out on a large dataset the first time someone requests it, but the  bucket creation and data transfer still complete in the background. The next time you try, it will just return the name of the existing bucket.",
      "votes": null
    },
    {
      "id": "810906",
      "postDate": "04/17/2020 12:03:03",
      "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> It's good news because currently I'm struggling with memory issue.\nNow most of the kernel load the text from dataframe and convert them to numpy array token on local memory. In my understanding, with this pipeline even if I use <code>GCS_DS_PATH = KaggleDatasets().get_gcs_path()</code>, still it's allocated to local memory. To avoid this, do we need to convert to token and save them beforehand, or can we convert text to token online with map function or something like that?</p>",
      "rawMarkdown": "mgornergoogle It's good news because currently I'm struggling with memory issue.\nNow most of the kernel load the text from dataframe and convert them to numpy array token on local memory. In my understanding, with this pipeline even if I use `GCS_DS_PATH = KaggleDatasets().get_gcs_path()`, still it's allocated to local memory. To avoid this, do we need to convert to token and save them beforehand, or can we convert text to token online with map function or something like that?",
      "votes": null
    },
    {
      "id": "814662",
      "postDate": "04/20/2020 21:13:09",
      "content": "<p>In short, you can use Kaggle data in 2 different locations:\n- locally: the Kaggle dataset attached to your notebook\n- on GCS: call <code>KaggleDatasets().get_gcs_path()</code> and Kaggle will happily move the data to a GCS bucket.</p>\n\n<p>What you do with the data after that is on you.</p>\n\n<p>However, if your dataset does not fit in memory and you want to load it continuously during training on TPU, the only way is to use the tf.data.Datasdet to load the data from GCS. It has to be GCS for this use case because the tf.data.Dataset pipeline will be executed by the TPU directly and there is single mass storage service the TPU can read from: GCS.</p>",
      "rawMarkdown": "In short, you can use Kaggle data in 2 different locations:\n- locally: the Kaggle dataset attached to your notebook\n- on GCS: call `KaggleDatasets().get_gcs_path()` and Kaggle will happily move the data to a GCS bucket.\n\nWhat you do with the data after that is on you.\n\nHowever, if your dataset does not fit in memory and you want to load it continuously during training on TPU, the only way is to use the tf.data.Datasdet to load the data from GCS. It has to be GCS for this use case because the tf.data.Dataset pipeline will be executed by the TPU directly and there is single mass storage service the TPU can read from: GCS.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 809455,
      "author_name": "adityaecdrid",
      "author_url": "",
      "post_date": "04/16/2020 07:56:59",
      "content": "<p>Because we are using data-set from local drives?</p>",
      "votes": null,
      "replies": [
        {
          "id": 809471,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "04/16/2020 08:09:21",
          "content": "<p>I am not sure, Aditya. </p>\n\n<p>In fact, it's a bit off-topic from this comp., but just for reference,\nI am playing around with TF2.1 image captioning for fun .. I copy all the codes from TF2.1 website, and luckily use Flickr / COCO dataset in locally located in Kaggle (instead of downloading using internet like in the example)</p>\n\n<p><a href=\"https://www.kaggle.com/ratthachat/image-captioning-using-attention/\">https://www.kaggle.com/ratthachat/image-captioning-using-attention/</a> (made public just for reference -- this kernel is still very messy)</p>\n\n<p>But in this case, it seems I have to use GCS_DS_PATH = KaggleDatasets().get_gcs_path() \n(even thought the data is in local drive too)</p>\n\n<p>And one more question, if somebody know :D ,  is there a limit on dataset size to use <code>get_gcs_path()</code> (e.g. I use full COCO2014 with 18GB and it got HTTP error)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 809479,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "04/16/2020 08:15:27",
          "content": "<p>Hmm interesting, IIRC, they move the data to GCS nearest to the VM we use automatically? (i don't remember where i read something similar to this;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 810149,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "04/16/2020 18:50:27",
      "content": "<p>If you training on data in memory, you can load the data any way you want.</p>\n\n<p>If you are using the tf.data.Dataset API to work with a dataset that does not fit in memory, then your data pipeline will be shipped to the TPU and executed there. The TPU will be doing the loading and the TPU can access data in GCS (Google Cloud Storage) only.</p>\n\n<p>Fortunately, Kaggle can get you a pointer to the competition data in GCS through:\n<code>\nGCS_DS_PATH = KaggleDatasets().get_gcs_path()\n</code>\nAs an added bonus, Kaggle will give you a GCS bucket that is always colocated with the TPU so you don't end up running in Europe and loading data from the west coast. That would be slow.</p>",
      "votes": null,
      "replies": [
        {
          "id": 810151,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "04/16/2020 18:53:11",
          "content": "<p>For data in memory, TPU training on numpy arrays directly works (in Keras model.fit()). Also, any memory array can be converted to a tf.data.Dataset with <code>Dataset.from_tensor_slices()</code> and then you can train on the tf.data.Dataset directly.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 810424,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "04/17/2020 00:05:50",
          "content": "<p>Thank you for explanation, Martin! <a href=\"/mgornergoogle\">@mgornergoogle</a> \nIs there any limitation on the size of the dataset to be used with <code>KaggleDatasets().get_gcs_path()</code> ?</p>\n\n<p>E.g. I tried a dataset of size 10-19GB and got an HTTP error. (internet is on)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 810508,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "04/17/2020 03:15:01",
          "content": "<p>Try again. Yes, in the current implementation it can time out on a large dataset the first time someone requests it, but the  bucket creation and data transfer still complete in the background. The next time you try, it will just return the name of the existing bucket.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 810906,
          "author_name": "bamps53",
          "author_url": "",
          "post_date": "04/17/2020 12:03:03",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> It's good news because currently I'm struggling with memory issue.\nNow most of the kernel load the text from dataframe and convert them to numpy array token on local memory. In my understanding, with this pipeline even if I use <code>GCS_DS_PATH = KaggleDatasets().get_gcs_path()</code>, still it's allocated to local memory. To avoid this, do we need to convert to token and save them beforehand, or can we convert text to token online with map function or something like that?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 814662,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "04/20/2020 21:13:09",
          "content": "<p>In short, you can use Kaggle data in 2 different locations:\n- locally: the Kaggle dataset attached to your notebook\n- on GCS: call <code>KaggleDatasets().get_gcs_path()</code> and Kaggle will happily move the data to a GCS bucket.</p>\n\n<p>What you do with the data after that is on you.</p>\n\n<p>However, if your dataset does not fit in memory and you want to load it continuously during training on TPU, the only way is to use the tf.data.Datasdet to load the data from GCS. It has to be GCS for this use case because the tf.data.Dataset pipeline will be executed by the TPU directly and there is single mass storage service the TPU can read from: GCS.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "809448": "Hi all, I am a newbie on TPU.\n\nI am studying @xhlulu great kernel and also other TPU kernels in image competitions.\nhttps://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\n\nIn competitions like Flower or Plant, it seems we need to use this command:\n`GCS_DS_PATH = KaggleDatasets().get_gcs_path()`\n\nI am a bit confused why we don't need it here.\n\nThanks!!",
    "809455": "Because we are using data-set from local drives?",
    "809471": "I am not sure, Aditya. \n\nIn fact, it's a bit off-topic from this comp., but just for reference,\nI am playing around with TF2.1 image captioning for fun .. I copy all the codes from TF2.1 website, and luckily use Flickr / COCO dataset in locally located in Kaggle (instead of downloading using internet like in the example)\n\nhttps://www.kaggle.com/ratthachat/image-captioning-using-attention/ (made public just for reference -- this kernel is still very messy)\n\nBut in this case, it seems I have to use GCS_DS_PATH = KaggleDatasets().get_gcs_path() \n(even thought the data is in local drive too)\n\nAnd one more question, if somebody know :D ,  is there a limit on dataset size to use `get_gcs_path()` (e.g. I use full COCO2014 with 18GB and it got HTTP error)",
    "809479": "Hmm interesting, IIRC, they move the data to GCS nearest to the VM we use automatically? (i don't remember where i read something similar to this;)",
    "810149": "If you training on data in memory, you can load the data any way you want.\n\nIf you are using the tf.data.Dataset API to work with a dataset that does not fit in memory, then your data pipeline will be shipped to the TPU and executed there. The TPU will be doing the loading and the TPU can access data in GCS (Google Cloud Storage) only.\n\nFortunately, Kaggle can get you a pointer to the competition data in GCS through:\n```\nGCS_DS_PATH = KaggleDatasets().get_gcs_path()\n```\nAs an added bonus, Kaggle will give you a GCS bucket that is always colocated with the TPU so you don't end up running in Europe and loading data from the west coast. That would be slow.",
    "810151": "For data in memory, TPU training on numpy arrays directly works (in Keras model.fit()). Also, any memory array can be converted to a tf.data.Dataset with `Dataset.from_tensor_slices()` and then you can train on the tf.data.Dataset directly.",
    "810424": "Thank you for explanation, Martin! @mgornergoogle \nIs there any limitation on the size of the dataset to be used with `KaggleDatasets().get_gcs_path()` ?\n\nE.g. I tried a dataset of size 10-19GB and got an HTTP error. (internet is on)",
    "810508": "Try again. Yes, in the current implementation it can time out on a large dataset the first time someone requests it, but the  bucket creation and data transfer still complete in the background. The next time you try, it will just return the name of the existing bucket.",
    "810906": "mgornergoogle It's good news because currently I'm struggling with memory issue.\nNow most of the kernel load the text from dataframe and convert them to numpy array token on local memory. In my understanding, with this pipeline even if I use `GCS_DS_PATH = KaggleDatasets().get_gcs_path()`, still it's allocated to local memory. To avoid this, do we need to convert to token and save them beforehand, or can we convert text to token online with map function or something like that?",
    "814662": "In short, you can use Kaggle data in 2 different locations:\n- locally: the Kaggle dataset attached to your notebook\n- on GCS: call `KaggleDatasets().get_gcs_path()` and Kaggle will happily move the data to a GCS bucket.\n\nWhat you do with the data after that is on you.\n\nHowever, if your dataset does not fit in memory and you want to load it continuously during training on TPU, the only way is to use the tf.data.Datasdet to load the data from GCS. It has to be GCS for this use case because the tf.data.Dataset pipeline will be executed by the TPU directly and there is single mass storage service the TPU can read from: GCS."
  },
  "source": "meta"
}