{
  "id": 170806,
  "title": "What is the efficient way to load data in colab?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/170806",
  "author_name": "lawrence",
  "post_date": "2020-07-29T06:16:49.463000",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I used kaggle API to download the data as in other competition, but tfrec and dcm files bothered me. I tried to access the images in .tfrec file, but I'm stucked where I don't know the format field of the tfrec file, and I couldn't read the content of it.</p>\n\n<p>Many others in discussion suggests download data from google cloud storage, but in order to do that I have to download the whole data first, then upload to the GCS bucket. It seems not to be an efficient way.</p>\n\n<p>How could I load the data in google colab efficiently? </p>",
  "messages": [
    {
      "id": 950034,
      "postDate": "2020-07-29T06:16:49.463Z",
      "content": "<p>I used kaggle API to download the data as in other competition, but tfrec and dcm files bothered me. I tried to access the images in .tfrec file, but I'm stucked where I don't know the format field of the tfrec file, and I couldn't read the content of it.</p>\n\n<p>Many others in discussion suggests download data from google cloud storage, but in order to do that I have to download the whole data first, then upload to the GCS bucket. It seems not to be an efficient way.</p>\n\n<p>How could I load the data in google colab efficiently? </p>",
      "rawMarkdown": "I used kaggle API to download the data as in other competition, but tfrec and dcm files bothered me. I tried to access the images in .tfrec file, but I'm stucked where I don't know the format field of the tfrec file, and I couldn't read the content of it.\n\nMany others in discussion suggests download data from google cloud storage, but in order to do that I have to download the whole data first, then upload to the GCS bucket. It seems not to be an efficient way.\n\nHow could I load the data in google colab efficiently? \n\n",
      "votes": 2
    },
    {
      "id": 952754,
      "postDate": "2020-07-31T07:31:47.207Z",
      "content": "<p>Hi there! Several notebooks I found so far only used the 'image' and 'target' feature from the provided .tfrec files. The format are as follow:\n<code>tfrecord_format = {\n        \"image\": tf.io.FixedLenFeature([], tf.string),\n        \"target\": tf.io.FixedLenFeature([], tf.int64)\n    }</code></p>\n\n<p>You can read the list of filenames as TFRecordDataset using\n<code>dataset = tf.data.TFRecordDataset(filenames)</code></p>\n\n<p>And then parse the \"examples\" using:\n<code>example = tf.io.parse_single_example(dataset, tfrecord_format)</code></p>\n\n<p>But this is still in string64 format, so to view the image you need to convert it back to the original format (fp32 maybe). I still don't know how to do that, but the others just use this parsed \"example\" and feed it to their training with tf.data API (with some preprocessing of course)</p>\n\n<p>I think you only need GCS bucket to run training on TPUs</p>",
      "rawMarkdown": "Hi there! Several notebooks I found so far only used the 'image' and 'target' feature from the provided .tfrec files. The format are as follow:\n`tfrecord_format = {\n        \"image\": tf.io.FixedLenFeature([], tf.string),\n        \"target\": tf.io.FixedLenFeature([], tf.int64)\n    }`\n\nYou can read the list of filenames as TFRecordDataset using\n`dataset = tf.data.TFRecordDataset(filenames)`\n\nAnd then parse the \"examples\" using:\n`example = tf.io.parse_single_example(dataset, tfrecord_format)`\n\nBut this is still in string64 format, so to view the image you need to convert it back to the original format (fp32 maybe). I still don't know how to do that, but the others just use this parsed \"example\" and feed it to their training with tf.data API (with some preprocessing of course)\n\nI think you only need GCS bucket to run training on TPUs"
    },
    {
      "id": 952097,
      "postDate": "2020-07-30T16:07:13.930Z",
      "content": "<p>If you are willing to spend like $10 or willing to get free GCP credits, your life will be a lot easier if you host the data on your own GCS bucket. That way you don't have to lookup the GCS_PATH every time you open your colab notebook. </p>",
      "rawMarkdown": "If you are willing to spend like $10 or willing to get free GCP credits, your life will be a lot easier if you host the data on your own GCS bucket. That way you don't have to lookup the GCS_PATH every time you open your colab notebook. "
    },
    {
      "id": 950506,
      "postDate": "2020-07-29T12:43:35.197Z",
      "content": "<p>Hey <a href=\"/laurencelin\">@laurencelin</a>  , \nif you are familier with TPU on colab and tfrecords, then you can directly use GCS_PATH  to directly load your data into functioning.\nWhereas , if you still want to work with jpeg images, then <a href=\"https://www.kaggle.com/prashantarorat/images-siim-512x512\">here</a> i have made a dataset with all jpeg images converted to 512x512x3 shape , which you can easily import into google colab using kaggle api .</p>",
      "rawMarkdown": "Hey @laurencelin  , \nif you are familier with TPU on colab and tfrecords, then you can directly use GCS_PATH  to directly load your data into functioning.\nWhereas , if you still want to work with jpeg images, then [here](https://www.kaggle.com/prashantarorat/images-siim-512x512) i have made a dataset with all jpeg images converted to 512x512x3 shape , which you can easily import into google colab using kaggle api ."
    },
    {
      "id": 950163,
      "postDate": "2020-07-29T08:10:21.973Z",
      "content": "<p>Hi, </p>\n\n<p>there is one public kernel with paths to GCS for mostly available public data. You can simply use those paths to construct your data. For example, in Triple Stratified KFold kernel from <a href=\"/cdeotte\">@cdeotte</a>, he has created a variable <code>GCS_PATH</code> which holds a path to train and test tfrecords, you simply can copy-paste the values of the variable in colab with a same or different variable. Sorry, I could not show you the code or example of it but it is a simple copy-paste of GCS path from Kaggle to colab and that's it. Your colab TPU will be able to access the data and you will be able to train models. One point to note is that Kaggle has TPU-V3 while colab has TPU-V2, so you need to use smaller batch size to run on colab as TPU-V2 has smaller RAM. </p>\n\n<p>I hope this will help. </p>",
      "rawMarkdown": "Hi, \n\nthere is one public kernel with paths to GCS for mostly available public data. You can simply use those paths to construct your data. For example, in Triple Stratified KFold kernel from @cdeotte, he has created a variable `GCS_PATH` which holds a path to train and test tfrecords, you simply can copy-paste the values of the variable in colab with a same or different variable. Sorry, I could not show you the code or example of it but it is a simple copy-paste of GCS path from Kaggle to colab and that's it. Your colab TPU will be able to access the data and you will be able to train models. One point to note is that Kaggle has TPU-V3 while colab has TPU-V2, so you need to use smaller batch size to run on colab as TPU-V2 has smaller RAM. \n\nI hope this will help. "
    },
    {
      "id": 951626,
      "postDate": "2020-07-30T09:19:47.367Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 952754,
      "author_name": "Sandhi Wangiyana",
      "author_url": "",
      "post_date": "2020-07-31T07:31:47.207000",
      "content": "<p>Hi there! Several notebooks I found so far only used the 'image' and 'target' feature from the provided .tfrec files. The format are as follow:\n<code>tfrecord_format = {\n        \"image\": tf.io.FixedLenFeature([], tf.string),\n        \"target\": tf.io.FixedLenFeature([], tf.int64)\n    }</code></p>\n\n<p>You can read the list of filenames as TFRecordDataset using\n<code>dataset = tf.data.TFRecordDataset(filenames)</code></p>\n\n<p>And then parse the \"examples\" using:\n<code>example = tf.io.parse_single_example(dataset, tfrecord_format)</code></p>\n\n<p>But this is still in string64 format, so to view the image you need to convert it back to the original format (fp32 maybe). I still don't know how to do that, but the others just use this parsed \"example\" and feed it to their training with tf.data API (with some preprocessing of course)</p>\n\n<p>I think you only need GCS bucket to run training on TPUs</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 952097,
      "author_name": "hongy",
      "author_url": "",
      "post_date": "2020-07-30T16:07:13.930000",
      "content": "<p>If you are willing to spend like $10 or willing to get free GCP credits, your life will be a lot easier if you host the data on your own GCS bucket. That way you don't have to lookup the GCS_PATH every time you open your colab notebook. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 950506,
      "author_name": "Prashant Arora",
      "author_url": "",
      "post_date": "2020-07-29T12:43:35.197000",
      "content": "<p>Hey <a href=\"/laurencelin\">@laurencelin</a>  , \nif you are familier with TPU on colab and tfrecords, then you can directly use GCS_PATH  to directly load your data into functioning.\nWhereas , if you still want to work with jpeg images, then <a href=\"https://www.kaggle.com/prashantarorat/images-siim-512x512\">here</a> i have made a dataset with all jpeg images converted to 512x512x3 shape , which you can easily import into google colab using kaggle api .</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 950163,
      "author_name": "Urvish",
      "author_url": "",
      "post_date": "2020-07-29T08:10:21.973000",
      "content": "<p>Hi, </p>\n\n<p>there is one public kernel with paths to GCS for mostly available public data. You can simply use those paths to construct your data. For example, in Triple Stratified KFold kernel from <a href=\"/cdeotte\">@cdeotte</a>, he has created a variable <code>GCS_PATH</code> which holds a path to train and test tfrecords, you simply can copy-paste the values of the variable in colab with a same or different variable. Sorry, I could not show you the code or example of it but it is a simple copy-paste of GCS path from Kaggle to colab and that's it. Your colab TPU will be able to access the data and you will be able to train models. One point to note is that Kaggle has TPU-V3 while colab has TPU-V2, so you need to use smaller batch size to run on colab as TPU-V2 has smaller RAM. </p>\n\n<p>I hope this will help. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 951626,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-30T09:19:47.367000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "950034": "I used kaggle API to download the data as in other competition, but tfrec and dcm files bothered me. I tried to access the images in .tfrec file, but I'm stucked where I don't know the format field of the tfrec file, and I couldn't read the content of it.\n\nMany others in discussion suggests download data from google cloud storage, but in order to do that I have to download the whole data first, then upload to the GCS bucket. It seems not to be an efficient way.\n\nHow could I load the data in google colab efficiently? \n\n",
    "952754": "Hi there! Several notebooks I found so far only used the 'image' and 'target' feature from the provided .tfrec files. The format are as follow:\n`tfrecord_format = {\n        \"image\": tf.io.FixedLenFeature([], tf.string),\n        \"target\": tf.io.FixedLenFeature([], tf.int64)\n    }`\n\nYou can read the list of filenames as TFRecordDataset using\n`dataset = tf.data.TFRecordDataset(filenames)`\n\nAnd then parse the \"examples\" using:\n`example = tf.io.parse_single_example(dataset, tfrecord_format)`\n\nBut this is still in string64 format, so to view the image you need to convert it back to the original format (fp32 maybe). I still don't know how to do that, but the others just use this parsed \"example\" and feed it to their training with tf.data API (with some preprocessing of course)\n\nI think you only need GCS bucket to run training on TPUs",
    "952097": "If you are willing to spend like $10 or willing to get free GCP credits, your life will be a lot easier if you host the data on your own GCS bucket. That way you don't have to lookup the GCS_PATH every time you open your colab notebook. ",
    "950506": "Hey @laurencelin  , \nif you are familier with TPU on colab and tfrecords, then you can directly use GCS_PATH  to directly load your data into functioning.\nWhereas , if you still want to work with jpeg images, then [here](https://www.kaggle.com/prashantarorat/images-siim-512x512) i have made a dataset with all jpeg images converted to 512x512x3 shape , which you can easily import into google colab using kaggle api .",
    "950163": "Hi, \n\nthere is one public kernel with paths to GCS for mostly available public data. You can simply use those paths to construct your data. For example, in Triple Stratified KFold kernel from @cdeotte, he has created a variable `GCS_PATH` which holds a path to train and test tfrecords, you simply can copy-paste the values of the variable in colab with a same or different variable. Sorry, I could not show you the code or example of it but it is a simple copy-paste of GCS path from Kaggle to colab and that's it. Your colab TPU will be able to access the data and you will be able to train models. One point to note is that Kaggle has TPU-V3 while colab has TPU-V2, so you need to use smaller batch size to run on colab as TPU-V2 has smaller RAM. \n\nI hope this will help. ",
    "951626": ""
  }
}