{
  "id": 232906,
  "title": "Using TFRecord for Easier Data Manipulation",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/232906",
  "author_name": "ryanl",
  "post_date": "2021-04-16T02:11:13.987000",
  "votes": 0,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I have found data i/o issues to be a common problem when optimizing my code. One common issue is that I go and build data processing pipelines in a single threaded architecture, then when trying to scale that in tensorflow I hit roadblocks. One way to solve that problem is to keep all your data read operations in native tf, especially using the new tf.data format, which allows you to easily use map + other data i/o optimizations (such as batch, prefetch and cache).  </p>\n<p>Towards this goal of making my data read operations use the tf.data format, I have created a script to segment cell line images and deposit the individual cells into TFRecords (which is a serialized data format created by google). I then show a quick example how to read the data back out. </p>\n<p>Hope you enjoy: </p>\n<p><a href=\"https://www.kaggle.com/ryanlstevens/process-hpa-files-write-read-tfrecord-data\" target=\"_blank\">https://www.kaggle.com/ryanlstevens/process-hpa-files-write-read-tfrecord-data</a></p>",
  "messages": [
    {
      "id": 1275214,
      "postDate": "2021-04-16T04:58:42.270Z",
      "content": "<p>That's quite useful , thanks! However,  I think the notebook hasn't run properly- I see <br>\n<code>Your notebook was stopped because it was idle for too long.</code><br>\nwhen I open then notebook link you shared.</p>",
      "rawMarkdown": "That's quite useful , thanks! However,  I think the notebook hasn't run properly- I see \n`Your notebook was stopped because it was idle for too long.`\nwhen I open then notebook link you shared.",
      "votes": 1,
      "replies": [
        {
          "id": 1275505,
          "postDate": "2021-04-16T11:32:05.837Z",
          "content": "<p>Thanks, just recreated it and uploaded a new version, should be fine now. </p>",
          "rawMarkdown": "Thanks, just recreated it and uploaded a new version, should be fine now. "
        }
      ]
    },
    {
      "id": 1275149,
      "postDate": "2021-04-16T02:11:13.987Z",
      "content": "<p>I have found data i/o issues to be a common problem when optimizing my code. One common issue is that I go and build data processing pipelines in a single threaded architecture, then when trying to scale that in tensorflow I hit roadblocks. One way to solve that problem is to keep all your data read operations in native tf, especially using the new tf.data format, which allows you to easily use map + other data i/o optimizations (such as batch, prefetch and cache).  </p>\n<p>Towards this goal of making my data read operations use the tf.data format, I have created a script to segment cell line images and deposit the individual cells into TFRecords (which is a serialized data format created by google). I then show a quick example how to read the data back out. </p>\n<p>Hope you enjoy: </p>\n<p><a href=\"https://www.kaggle.com/ryanlstevens/process-hpa-files-write-read-tfrecord-data\" target=\"_blank\">https://www.kaggle.com/ryanlstevens/process-hpa-files-write-read-tfrecord-data</a></p>",
      "rawMarkdown": "I have found data i/o issues to be a common problem when optimizing my code. One common issue is that I go and build data processing pipelines in a single threaded architecture, then when trying to scale that in tensorflow I hit roadblocks. One way to solve that problem is to keep all your data read operations in native tf, especially using the new tf.data format, which allows you to easily use map + other data i/o optimizations (such as batch, prefetch and cache).  \n\nTowards this goal of making my data read operations use the tf.data format, I have created a script to segment cell line images and deposit the individual cells into TFRecords (which is a serialized data format created by google). I then show a quick example how to read the data back out. \n\nHope you enjoy: \n\nhttps://www.kaggle.com/ryanlstevens/process-hpa-files-write-read-tfrecord-data"
    }
  ],
  "comments": [
    {
      "id": 1275214,
      "author_name": "Satwik",
      "author_url": "",
      "post_date": "2021-04-16T04:58:42.270000",
      "content": "<p>That's quite useful , thanks! However,  I think the notebook hasn't run properly- I see <br>\n<code>Your notebook was stopped because it was idle for too long.</code><br>\nwhen I open then notebook link you shared.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1275505,
          "author_name": "ryanl",
          "author_url": "",
          "post_date": "2021-04-16T11:32:05.837000",
          "content": "<p>Thanks, just recreated it and uploaded a new version, should be fine now. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1275214": "That's quite useful , thanks! However,  I think the notebook hasn't run properly- I see \n`Your notebook was stopped because it was idle for too long.`\nwhen I open then notebook link you shared.",
    "1275149": "I have found data i/o issues to be a common problem when optimizing my code. One common issue is that I go and build data processing pipelines in a single threaded architecture, then when trying to scale that in tensorflow I hit roadblocks. One way to solve that problem is to keep all your data read operations in native tf, especially using the new tf.data format, which allows you to easily use map + other data i/o optimizations (such as batch, prefetch and cache).  \n\nTowards this goal of making my data read operations use the tf.data format, I have created a script to segment cell line images and deposit the individual cells into TFRecords (which is a serialized data format created by google). I then show a quick example how to read the data back out. \n\nHope you enjoy: \n\nhttps://www.kaggle.com/ryanlstevens/process-hpa-files-write-read-tfrecord-data"
  }
}