{
  "id": 39698,
  "title": "Question on Data Management",
  "url": "/competitions/cdiscount-image-classification-challenge/discussion/39698",
  "author_name": "",
  "post_date": "2017-09-19T14:44:50.705287400Z",
  "votes": 1,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Hi, this is the first time that I deal with this size of dataset . So far my strategy is not successful due to memory and disk limit. I was wondering if someone would like to share how to process the train data?</p>\n\n<p>My current strategy summarizes as follows:</p>\n\n<ol>\n<li><p>create a VM instances on Google Cloud</p></li>\n<li><p>download train.bson and test.bson data to the vm</p></li>\n<li><p>load bson data into python using Iversion's kernal at <a href=\"https://www.kaggle.com/inversion/processing-bson-files\">https://www.kaggle.com/inversion/processing-bson-files</a></p></li>\n<li><p>For each image, I was trying to reshape into 1d array and combine it with product_id and category_id. put them as an row in my numpy array.</p></li>\n<li><p>Due to the memory limit, my numpy array cannot hold entire training set. Instead, I was trying to slice the training set into chunks and save each chunk as a separate numpy array and store it into hdf5. However, the corresponding hdf5 can be way much larger than the bson file and even with 50,000 samples, it has been around 40 GB.</p></li>\n</ol>\n\n<p>Can someone help me with the data management and processing?</p>",
  "messages": [
    {
      "id": "222641",
      "postDate": "09/19/2017 14:44:50",
      "content": "<p>Hi, this is the first time that I deal with this size of dataset . So far my strategy is not successful due to memory and disk limit. I was wondering if someone would like to share how to process the train data?</p>\n\n<p>My current strategy summarizes as follows:</p>\n\n<ol>\n<li><p>create a VM instances on Google Cloud</p></li>\n<li><p>download train.bson and test.bson data to the vm</p></li>\n<li><p>load bson data into python using Iversion's kernal at <a href=\"https://www.kaggle.com/inversion/processing-bson-files\">https://www.kaggle.com/inversion/processing-bson-files</a></p></li>\n<li><p>For each image, I was trying to reshape into 1d array and combine it with product_id and category_id. put them as an row in my numpy array.</p></li>\n<li><p>Due to the memory limit, my numpy array cannot hold entire training set. Instead, I was trying to slice the training set into chunks and save each chunk as a separate numpy array and store it into hdf5. However, the corresponding hdf5 can be way much larger than the bson file and even with 50,000 samples, it has been around 40 GB.</p></li>\n</ol>\n\n<p>Can someone help me with the data management and processing?</p>",
      "rawMarkdown": "Hi, this is the first time that I deal with this size of dataset . So far my strategy is not successful due to memory and disk limit. I was wondering if someone would like to share how to process the train data?\n\nMy current strategy summarizes as follows:\n\n1. create a VM instances on Google Cloud\n\n2. download train.bson and test.bson data to the vm\n\n3. load bson data into python using Iversion's kernal at https://www.kaggle.com/inversion/processing-bson-files\n\n4. For each image, I was trying to reshape into 1d array and combine it with product_id and category_id. put them as an row in my numpy array.\n\n5. Due to the memory limit, my numpy array cannot hold entire training set. Instead, I was trying to slice the training set into chunks and save each chunk as a separate numpy array and store it into hdf5. However, the corresponding hdf5 can be way much larger than the bson file and even with 50,000 samples, it has been around 40 GB.\n\nCan someone help me with the data management and processing?",
      "votes": null
    },
    {
      "id": "222671",
      "postDate": "09/19/2017 16:45:46",
      "content": "<p>I think dealing with the huge amount of data is as much a challenge as doing the actual classification. ;-)</p>",
      "rawMarkdown": "I think dealing with the huge amount of data is as much a challenge as doing the actual classification. ;-)",
      "votes": null
    },
    {
      "id": "222697",
      "postDate": "09/19/2017 18:03:39",
      "content": "<p>vfdev5 spared me some work to the approach I wanted with his kernel: <a href=\"https://www.kaggle.com/vfdev5/random-item-access\">https://www.kaggle.com/vfdev5/random-item-access</a> </p>\n\n<p>Did you ever had to work with data image augmentation? Think about it in the same way, but here you have a single file (much more manageable also, 60Gb) in which you can come pick a fixed number of samples (dependant of your memory) as much as you want (think of it as epochs), acting like a generator.</p>",
      "rawMarkdown": "vfdev5 spared me some work to the approach I wanted with his kernel: https://www.kaggle.com/vfdev5/random-item-access \n\nDid you ever had to work with data image augmentation? Think about it in the same way, but here you have a single file (much more manageable also, 60Gb) in which you can come pick a fixed number of samples (dependant of your memory) as much as you want (think of it as epochs), acting like a generator.",
      "votes": null
    },
    {
      "id": "223853",
      "postDate": "09/23/2017 22:00:30",
      "content": "<p>Are you using Keras? If so, have you tried fit_generator? It's a way to train in shunks...</p>",
      "rawMarkdown": "Are you using Keras? If so, have you tried fit_generator? It's a way to train in shunks...",
      "votes": null
    },
    {
      "id": "224039",
      "postDate": "09/24/2017 21:10:15",
      "content": "<p>If you dump data into compressed tfrecord files and work with tensorflow you can get down with train dataset to 47GB.</p>",
      "rawMarkdown": "If you dump data into compressed tfrecord files and work with tensorflow you can get down with train dataset to 47GB.",
      "votes": null
    },
    {
      "id": "224292",
      "postDate": "09/25/2017 20:27:22",
      "content": "<p>Hi Marcin, thanks very much for your post. I was trying to use your strategy and almost succeed. However, I keep got error from read the data from tfrecord files. I was trying to fix it for hours but had no luck. My code is as follows. Can you please take a look and help me with where I am wrong? Thanks</p>\n\n<p>The first block works fine. It reads the bson data and dump it to TFRecords</p>\n\n<pre><code>def _int64_feature(value):\n  return tf.train.Feature(int64_list=tf.train.Int64List(value=[value]))\n\ndef _bytes_feature(value):\n  return tf.train.Feature(bytes_list=tf.train.BytesList(value=[value]))\n\ntrain_filename = 'train.tfrecords'\nwriter = tf.python_io.TFRecordWriter(train_filename)\n\ndata = bson.decode_file_iter(open('../data/train_example.bson','rb'))\nwith open('../data/train_example_array.txt','ab') as f:\n    for c,d in enumerate(data):\n        product_id = d['_id']\n        category_id = d['category_id']\n        for e,pic in enumerate(d['imgs']): # loop through each picture \n            img = imread(io.BytesIO(pic['picture'])) # output is a 180*180*3=97200 array\n            label = category_id\n            print(img.shape)\n            feature = {'train/label': _int64_feature(label),\n                       'train/image': _bytes_feature(tf.compat.as_bytes(img.tostring()))\n            example = tf.train.Example(features=tf.train.Features(feature=feature)\n            writer.write(example.SerializeToString())\nwriter.close()\n</code></pre>\n\n<p>However, I get error from the following read from the following reading part. The error says</p>\n\n<p>INFO:tensorflow:Error reported to Coordinator: , Input to reshape is a tensor with 24300 values, but the requested shape has 97200\n         [[Node: Reshape = Reshape[T=DT_FLOAT, Tshape=DT_INT32, _device=\"/job:localhost/replica:0/task:0/cpu:0\"](DecodeRaw, Reshape/shape)]]</p>\n\n<p>I can only say the error produced from command:</p>\n\n<pre><code>threads = tf.train.start_queue_runners(coord=coord) \n</code></pre>\n\n<p>but I don't know why the input tensor is 24300 values. Even if I change the reshape command to (90*90*3=24300). It still gives me the same error. Thanks for help!!!</p>\n\n<pre><code># Read TFRecord File\nimport matplotlib.pyplot as plt\n\ndata_path = 'train.tfrecords'\n\nwith tf.Session() as sess:\n    feature = {'train/label': tf.FixedLenFeature([],tf.int64),\n               'train/image': tf.FixedLenFeature([],tf.string)}\n\n    filename_queue = tf.train.string_input_producer([data_path],num_epochs=1)\n\n    reader = tf.TFRecordReader()\n    _, serialized_example = reader.read(filename_queue)\n\n    features = tf.parse_single_example(serialized_example,features=feature)\n\n    image = tf.decode_raw(features['train/image'],tf.float32)\n\n    label = tf.cast(features['train/label'],tf.int32)\n\n    image = tf.reshape(image,[180,180,3])\n\n\n\n    images,labels = tf.train.shuffle_batch([image,label],batch_size=10,capacity=30,num_threads=1,min_after_dequeue=10)\n\n     # Initialize all global and local variables\n    init_op = tf.group(tf.global_variables_initializer(), tf.local_variables_initializer())\n    sess.run(init_op)\n\n\n\n    # Create a coordinator and run all QueueRunner objects\n\n\n    coord = tf.train.Coordinator()\n    threads = tf.train.start_queue_runners(coord=coord)\n    for batch_index in range(5):\n        img, lbl = sess.run([images, labels])\n        img = img.astype(np.uint8)\n        for j in range(6):\n            plt.subplot(2, 3, j+1)\n            plt.imshow(img[j, ...])\n            #plt.title('cat' if lbl[j]==0 else 'dog')\n        plt.show()\n\n    coord.request_stop()\n    sess.close()\n</code></pre>",
      "rawMarkdown": "Hi Marcin, thanks very much for your post. I was trying to use your strategy and almost succeed. However, I keep got error from read the data from tfrecord files. I was trying to fix it for hours but had no luck. My code is as follows. Can you please take a look and help me with where I am wrong? Thanks\n\nThe first block works fine. It reads the bson data and dump it to TFRecords\n\n    def _int64_feature(value):\n      return tf.train.Feature(int64_list=tf.train.Int64List(value=[value]))\n\n    def _bytes_feature(value):\n      return tf.train.Feature(bytes_list=tf.train.BytesList(value=[value]))\n\n    train_filename = 'train.tfrecords'\n    writer = tf.python_io.TFRecordWriter(train_filename)\n    \n    data = bson.decode_file_iter(open('../data/train_example.bson','rb'))\n    with open('../data/train_example_array.txt','ab') as f:\n        for c,d in enumerate(data):\n            product_id = d['_id']\n            category_id = d['category_id']\n            for e,pic in enumerate(d['imgs']): # loop through each picture \n                img = imread(io.BytesIO(pic['picture'])) # output is a 180*180*3=97200 array\n                label = category_id\n                print(img.shape)\n                feature = {'train/label': _int64_feature(label),\n                           'train/image': _bytes_feature(tf.compat.as_bytes(img.tostring()))\n                example = tf.train.Example(features=tf.train.Features(feature=feature)\n                writer.write(example.SerializeToString())\n    writer.close()\n    \nHowever, I get error from the following read from the following reading part. The error says\n\n INFO:tensorflow:Error reported to Coordinator:",
      "votes": null
    },
    {
      "id": "224294",
      "postDate": "09/25/2017 20:36:54",
      "content": "<p>1) Maybe have a look at the data types when you write and read the file. I see some descrepancies in using Int64 and Int32. </p>\n\n<p>2) Also when working with tensorflow it is a good practice to name all transofrmations, so debugger will point you straight to the part that causes error.</p>\n\n<p>3) If you use WITH, you od not need to explicitly close session. </p>\n\n<p>Have a looks at a sample kernel how to read created TFRecord files:\n<a href=\"https://www.kaggle.com/mpekalski/reading-tfrecord/\">https://www.kaggle.com/mpekalski/reading-tfrecord/</a></p>",
      "rawMarkdown": "1) Maybe have a look at the data types when you write and read the file. I see some descrepancies in using Int64 and Int32. \n\n2) Also when working with tensorflow it is a good practice to name all transofrmations, so debugger will point you straight to the part that causes error.\n\n3) If you use WITH, you od not need to explicitly close session. \n\nHave a looks at a sample kernel how to read created TFRecord files:\nhttps://www.kaggle.com/mpekalski/reading-tfrecord/",
      "votes": null
    },
    {
      "id": "224502",
      "postDate": "09/26/2017 15:54:19",
      "content": "<p>I finally find a way to get around that error although still confused at where I was wrong. Thanks for sharing your trick. </p>",
      "rawMarkdown": "I finally find a way to get around that error although still confused at where I was wrong. Thanks for sharing your trick.",
      "votes": null
    },
    {
      "id": "224624",
      "postDate": "09/27/2017 03:21:30",
      "content": "<p>Hi Marcin, how you are able to put the entire training data into 47GB compressed TFRecords? I tested my code and found even with 500,000 images it has already taken up 31GB.</p>\n\n<p>my code is as follows:</p>\n\n<pre><code>import numpy as np\nimport io\nimport bson\nimport multiprocessing as mp\nimport sys\nimport tensorflow as tf\nfrom skimage.data import imread\n\ndef _int64_feature(value):\n  return tf.train.Feature(int64_list=tf.train.Int64List(value=[value]))\ndef _bytes_feature(value):\n  return tf.train.Feature(bytes_list=tf.train.BytesList(value=[value]))\ntrain_filename = '../data/train.tfrecords'\nopts = tf.python_io.TFRecordOptions(tf.python_io.TFRecordCompressionType.ZLIB)\nwriter = tf.python_io.TFRecordWriter(train_filename,options=opts)\ndata = bson.decode_file_iter(open('../data/train.bson','rb'))\nfor c,d in enumerate(data):\n    product_id = d['_id']\n    category_id = d['category_id']\n    for e,pic in enumerate(d['imgs']): # loop through each picture \n        img = imread(io.BytesIO(pic['picture'])) # output is a 180*180*3=97200 array\n        label = category_id\n        feature = {'label': _int64_feature(label),\n                   'image': _bytes_feature(img.tostring())}\n        # Create an example protocol buffer\n        example = tf.train.Example(features=tf.train.Features(feature=feature))\n        # Serialize to string and write on the file\n        writer.write(example.SerializeToString())\n    if c % 1000 == 0:\n        print(c)\nwriter.close()\n</code></pre>",
      "rawMarkdown": "Hi Marcin, how you are able to put the entire training data into 47GB compressed TFRecords? I tested my code and found even with 500,000 images it has already taken up 31GB.\n\nmy code is as follows:\n\n    import numpy as np\n    import io\n    import bson\n    import multiprocessing as mp\n    import sys\n    import tensorflow as tf\n    from skimage.data import imread\n   \n    def _int64_feature(value):\n      return tf.train.Feature(int64_list=tf.train.Int64List(value=[value]))\n    def _bytes_feature(value):\n      return tf.train.Feature(bytes_list=tf.train.BytesList(value=[value]))\n    train_filename = '../data/train.tfrecords'\n    opts = tf.python_io.TFRecordOptions(tf.python_io.TFRecordCompressionType.ZLIB)\n    writer = tf.python_io.TFRecordWriter(train_filename,options=opts)\n    data = bson.decode_file_iter(open('../data/train.bson','rb'))\n    for c,d in enumerate(data):\n        product_id = d['_id']\n        category_id = d['category_id']\n        for e,pic in enumerate(d['imgs']): # loop through each picture \n            img = imread(io.BytesIO(pic['picture'])) # output is a 180*180*3=97200 array\n            label = category_id\n            feature = {'label': _int64_feature(label),\n                       'image': _bytes_feature(img.tostring())}\n            # Create an example protocol buffer\n            example = tf.train.Example(features=tf.train.Features(feature=feature))\n            # Serialize to string and write on the file\n            writer.write(example.SerializeToString())\n        if c % 1000 == 0:\n            print(c)\n    writer.close()",
      "votes": null
    },
    {
      "id": "224650",
      "postDate": "09/27/2017 05:43:34",
      "content": "<p>@Wenbo, my code is in kernels <a href=\"https://www.kaggle.com/mpekalski/convert-bson-to-tfrecord\">https://www.kaggle.com/mpekalski/convert-bson-to-tfrecord</a></p>\n\n<p>From what I see you are decoding the image instead of storing it raw as you read it from BSON file.</p>",
      "rawMarkdown": "Wenbo, my code is in kernels https://www.kaggle.com/mpekalski/convert-bson-to-tfrecord\n\nFrom what I see you are decoding the image instead of storing it raw as you read it from BSON file.",
      "votes": null
    },
    {
      "id": "225280",
      "postDate": "09/28/2017 16:37:54",
      "content": "<p>@Marcin, instead of decoding, I was trying to store the bytes file to TFRecords. However, when I was trying to read from TFRecords, I got dimension error again. someone told me to add </p>\n\n<pre><code>img_raw = skimage.io.imread(io.StringIO(img_raw).astype(np.uint8).tostring())\n</code></pre>\n\n<p>before write img_raw to disk (img_raw is pic['picture'] which i guess is already bytes file). I don't get the logic why I need to add that line.</p>\n\n<p>My code and error is at:\n<a href=\"https://github.com/wenbo5565/admin/blob/master/Sample%2BRead%2BCompressed%2BTFRecord.ipynb\">https://github.com/wenbo5565/admin/blob/master/Sample%2BRead%2BCompressed%2BTFRecord.ipynb</a></p>\n\n<p>Appreciate your help!</p>",
      "rawMarkdown": "Marcin, instead of decoding, I was trying to store the bytes file to TFRecords. However, when I was trying to read from TFRecords, I got dimension error again. someone told me to add \n\n    img_raw = skimage.io.imread(io.StringIO(img_raw).astype(np.uint8).tostring())\n\nbefore write img_raw to disk (img_raw is pic['picture'] which i guess is already bytes file). I don't get the logic why I need to add that line.\n\nMy code and error is at:\nhttps://github.com/wenbo5565/admin/blob/master/Sample%2BRead%2BCompressed%2BTFRecord.ipynb\n\nAppreciate your help!",
      "votes": null
    },
    {
      "id": "225680",
      "postDate": "09/29/2017 19:49:27",
      "content": "<p>Sorry @Wenbo, but I have very limited amount of time I can spend on Kaggle and I cannot spare it at the moment to debug your code. </p>",
      "rawMarkdown": "Sorry @Wenbo, but I have very limited amount of time I can spend on Kaggle and I cannot spare it at the moment to debug your code.",
      "votes": null
    },
    {
      "id": "225717",
      "postDate": "09/29/2017 22:09:50",
      "content": "<p>@Marcin, I already fixed it up. Thanks for all the help so far.</p>",
      "rawMarkdown": "Marcin, I already fixed it up. Thanks for all the help so far.",
      "votes": null
    },
    {
      "id": "225893",
      "postDate": "09/30/2017 09:11:05",
      "content": "<p>I did not manage to create any model working with it yet, but in general tensorflow provides means for processing images. After reading stream of bytes you should use <strong>tf.image.decode_jpeg</strong> that returns a <strong>tf.uint8</strong> tensor but we want float32 as input to the model, so we have to convert it.</p>\n\n<pre><code>features =  {\n             'height': tf.FixedLenFeature((1), tf.int64),\n             'width': tf.FixedLenFeature((1), tf.int64),\n             'img_raw': tf.FixedLenFeature((), tf.string),\n             'category_id': tf.FixedLenFeature((1), tf.int64),\n             'product_id': tf.FixedLenFeature((1), tf.int64)                 \n            }\nparsed_features = tf.parse_single_example(record, features, name=\"parse_single_example\")\nimg = tf.image.decode_jpeg(parsed_features[\"img_raw\"], channels=3)\nimg.set_shape([180,180,3])\nimg = tf.image.convert_image_dtype(img, tf.float32)\n</code></pre>\n\n<p>then plotting image like that with matplotlib would simply be</p>\n\n<pre><code>with tf.Session() as sess:\n    tmp = sess.run(img)\n    plt.imshow(tmp)\n</code></pre>",
      "rawMarkdown": "I did not manage to create any model working with it yet, but in general tensorflow provides means for processing images. After reading stream of bytes you should use **tf.image.decode_jpeg** that returns a **tf.uint8** tensor but we want float32 as input to the model, so we have to convert it.\n\n\n    features =  {\n                 'height': tf.FixedLenFeature((1), tf.int64),\n                 'width': tf.FixedLenFeature((1), tf.int64),\n                 'img_raw': tf.FixedLenFeature((), tf.string),\n                 'category_id': tf.FixedLenFeature((1), tf.int64),\n                 'product_id': tf.FixedLenFeature((1), tf.int64)                 \n                }\n    parsed_features = tf.parse_single_example(record, features, name=\"parse_single_example\")\n    img = tf.image.decode_jpeg(parsed_features[\"img_raw\"], channels=3)\n    img.set_shape([180,180,3])\n    img = tf.image.convert_image_dtype(img, tf.float32)\n\nthen plotting image like that with matplotlib would simply be\n    \n    with tf.Session() as sess:\n        tmp = sess.run(img)\n        plt.imshow(tmp)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 222671,
      "author_name": "humananalog",
      "author_url": "",
      "post_date": "09/19/2017 16:45:46",
      "content": "<p>I think dealing with the huge amount of data is as much a challenge as doing the actual classification. ;-)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 222697,
      "author_name": "gaarv1911",
      "author_url": "",
      "post_date": "09/19/2017 18:03:39",
      "content": "<p>vfdev5 spared me some work to the approach I wanted with his kernel: <a href=\"https://www.kaggle.com/vfdev5/random-item-access\">https://www.kaggle.com/vfdev5/random-item-access</a> </p>\n\n<p>Did you ever had to work with data image augmentation? Think about it in the same way, but here you have a single file (much more manageable also, 60Gb) in which you can come pick a fixed number of samples (dependant of your memory) as much as you want (think of it as epochs), acting like a generator.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 223853,
      "author_name": "aloisiodn",
      "author_url": "",
      "post_date": "09/23/2017 22:00:30",
      "content": "<p>Are you using Keras? If so, have you tried fit_generator? It's a way to train in shunks...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 224039,
      "author_name": "mpekalski",
      "author_url": "",
      "post_date": "09/24/2017 21:10:15",
      "content": "<p>If you dump data into compressed tfrecord files and work with tensorflow you can get down with train dataset to 47GB.</p>",
      "votes": null,
      "replies": [
        {
          "id": 224292,
          "author_name": "wenbo5565",
          "author_url": "",
          "post_date": "09/25/2017 20:27:22",
          "content": "<p>Hi Marcin, thanks very much for your post. I was trying to use your strategy and almost succeed. However, I keep got error from read the data from tfrecord files. I was trying to fix it for hours but had no luck. My code is as follows. Can you please take a look and help me with where I am wrong? Thanks</p>\n\n<p>The first block works fine. It reads the bson data and dump it to TFRecords</p>\n\n<pre><code>def _int64_feature(value):\n  return tf.train.Feature(int64_list=tf.train.Int64List(value=[value]))\n\ndef _bytes_feature(value):\n  return tf.train.Feature(bytes_list=tf.train.BytesList(value=[value]))\n\ntrain_filename = 'train.tfrecords'\nwriter = tf.python_io.TFRecordWriter(train_filename)\n\ndata = bson.decode_file_iter(open('../data/train_example.bson','rb'))\nwith open('../data/train_example_array.txt','ab') as f:\n    for c,d in enumerate(data):\n        product_id = d['_id']\n        category_id = d['category_id']\n        for e,pic in enumerate(d['imgs']): # loop through each picture \n            img = imread(io.BytesIO(pic['picture'])) # output is a 180*180*3=97200 array\n            label = category_id\n            print(img.shape)\n            feature = {'train/label': _int64_feature(label),\n                       'train/image': _bytes_feature(tf.compat.as_bytes(img.tostring()))\n            example = tf.train.Example(features=tf.train.Features(feature=feature)\n            writer.write(example.SerializeToString())\nwriter.close()\n</code></pre>\n\n<p>However, I get error from the following read from the following reading part. The error says</p>\n\n<p>INFO:tensorflow:Error reported to Coordinator: , Input to reshape is a tensor with 24300 values, but the requested shape has 97200\n         [[Node: Reshape = Reshape[T=DT_FLOAT, Tshape=DT_INT32, _device=\"/job:localhost/replica:0/task:0/cpu:0\"](DecodeRaw, Reshape/shape)]]</p>\n\n<p>I can only say the error produced from command:</p>\n\n<pre><code>threads = tf.train.start_queue_runners(coord=coord) \n</code></pre>\n\n<p>but I don't know why the input tensor is 24300 values. Even if I change the reshape command to (90*90*3=24300). It still gives me the same error. Thanks for help!!!</p>\n\n<pre><code># Read TFRecord File\nimport matplotlib.pyplot as plt\n\ndata_path = 'train.tfrecords'\n\nwith tf.Session() as sess:\n    feature = {'train/label': tf.FixedLenFeature([],tf.int64),\n               'train/image': tf.FixedLenFeature([],tf.string)}\n\n    filename_queue = tf.train.string_input_producer([data_path],num_epochs=1)\n\n    reader = tf.TFRecordReader()\n    _, serialized_example = reader.read(filename_queue)\n\n    features = tf.parse_single_example(serialized_example,features=feature)\n\n    image = tf.decode_raw(features['train/image'],tf.float32)\n\n    label = tf.cast(features['train/label'],tf.int32)\n\n    image = tf.reshape(image,[180,180,3])\n\n\n\n    images,labels = tf.train.shuffle_batch([image,label],batch_size=10,capacity=30,num_threads=1,min_after_dequeue=10)\n\n     # Initialize all global and local variables\n    init_op = tf.group(tf.global_variables_initializer(), tf.local_variables_initializer())\n    sess.run(init_op)\n\n\n\n    # Create a coordinator and run all QueueRunner objects\n\n\n    coord = tf.train.Coordinator()\n    threads = tf.train.start_queue_runners(coord=coord)\n    for batch_index in range(5):\n        img, lbl = sess.run([images, labels])\n        img = img.astype(np.uint8)\n        for j in range(6):\n            plt.subplot(2, 3, j+1)\n            plt.imshow(img[j, ...])\n            #plt.title('cat' if lbl[j]==0 else 'dog')\n        plt.show()\n\n    coord.request_stop()\n    sess.close()\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 224294,
          "author_name": "mpekalski",
          "author_url": "",
          "post_date": "09/25/2017 20:36:54",
          "content": "<p>1) Maybe have a look at the data types when you write and read the file. I see some descrepancies in using Int64 and Int32. </p>\n\n<p>2) Also when working with tensorflow it is a good practice to name all transofrmations, so debugger will point you straight to the part that causes error.</p>\n\n<p>3) If you use WITH, you od not need to explicitly close session. </p>\n\n<p>Have a looks at a sample kernel how to read created TFRecord files:\n<a href=\"https://www.kaggle.com/mpekalski/reading-tfrecord/\">https://www.kaggle.com/mpekalski/reading-tfrecord/</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 224502,
          "author_name": "wenbo5565",
          "author_url": "",
          "post_date": "09/26/2017 15:54:19",
          "content": "<p>I finally find a way to get around that error although still confused at where I was wrong. Thanks for sharing your trick. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 224624,
          "author_name": "wenbo5565",
          "author_url": "",
          "post_date": "09/27/2017 03:21:30",
          "content": "<p>Hi Marcin, how you are able to put the entire training data into 47GB compressed TFRecords? I tested my code and found even with 500,000 images it has already taken up 31GB.</p>\n\n<p>my code is as follows:</p>\n\n<pre><code>import numpy as np\nimport io\nimport bson\nimport multiprocessing as mp\nimport sys\nimport tensorflow as tf\nfrom skimage.data import imread\n\ndef _int64_feature(value):\n  return tf.train.Feature(int64_list=tf.train.Int64List(value=[value]))\ndef _bytes_feature(value):\n  return tf.train.Feature(bytes_list=tf.train.BytesList(value=[value]))\ntrain_filename = '../data/train.tfrecords'\nopts = tf.python_io.TFRecordOptions(tf.python_io.TFRecordCompressionType.ZLIB)\nwriter = tf.python_io.TFRecordWriter(train_filename,options=opts)\ndata = bson.decode_file_iter(open('../data/train.bson','rb'))\nfor c,d in enumerate(data):\n    product_id = d['_id']\n    category_id = d['category_id']\n    for e,pic in enumerate(d['imgs']): # loop through each picture \n        img = imread(io.BytesIO(pic['picture'])) # output is a 180*180*3=97200 array\n        label = category_id\n        feature = {'label': _int64_feature(label),\n                   'image': _bytes_feature(img.tostring())}\n        # Create an example protocol buffer\n        example = tf.train.Example(features=tf.train.Features(feature=feature))\n        # Serialize to string and write on the file\n        writer.write(example.SerializeToString())\n    if c % 1000 == 0:\n        print(c)\nwriter.close()\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 224650,
          "author_name": "mpekalski",
          "author_url": "",
          "post_date": "09/27/2017 05:43:34",
          "content": "<p>@Wenbo, my code is in kernels <a href=\"https://www.kaggle.com/mpekalski/convert-bson-to-tfrecord\">https://www.kaggle.com/mpekalski/convert-bson-to-tfrecord</a></p>\n\n<p>From what I see you are decoding the image instead of storing it raw as you read it from BSON file.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 225280,
          "author_name": "wenbo5565",
          "author_url": "",
          "post_date": "09/28/2017 16:37:54",
          "content": "<p>@Marcin, instead of decoding, I was trying to store the bytes file to TFRecords. However, when I was trying to read from TFRecords, I got dimension error again. someone told me to add </p>\n\n<pre><code>img_raw = skimage.io.imread(io.StringIO(img_raw).astype(np.uint8).tostring())\n</code></pre>\n\n<p>before write img_raw to disk (img_raw is pic['picture'] which i guess is already bytes file). I don't get the logic why I need to add that line.</p>\n\n<p>My code and error is at:\n<a href=\"https://github.com/wenbo5565/admin/blob/master/Sample%2BRead%2BCompressed%2BTFRecord.ipynb\">https://github.com/wenbo5565/admin/blob/master/Sample%2BRead%2BCompressed%2BTFRecord.ipynb</a></p>\n\n<p>Appreciate your help!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 225680,
          "author_name": "mpekalski",
          "author_url": "",
          "post_date": "09/29/2017 19:49:27",
          "content": "<p>Sorry @Wenbo, but I have very limited amount of time I can spend on Kaggle and I cannot spare it at the moment to debug your code. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 225717,
          "author_name": "wenbo5565",
          "author_url": "",
          "post_date": "09/29/2017 22:09:50",
          "content": "<p>@Marcin, I already fixed it up. Thanks for all the help so far.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 225893,
          "author_name": "mpekalski",
          "author_url": "",
          "post_date": "09/30/2017 09:11:05",
          "content": "<p>I did not manage to create any model working with it yet, but in general tensorflow provides means for processing images. After reading stream of bytes you should use <strong>tf.image.decode_jpeg</strong> that returns a <strong>tf.uint8</strong> tensor but we want float32 as input to the model, so we have to convert it.</p>\n\n<pre><code>features =  {\n             'height': tf.FixedLenFeature((1), tf.int64),\n             'width': tf.FixedLenFeature((1), tf.int64),\n             'img_raw': tf.FixedLenFeature((), tf.string),\n             'category_id': tf.FixedLenFeature((1), tf.int64),\n             'product_id': tf.FixedLenFeature((1), tf.int64)                 \n            }\nparsed_features = tf.parse_single_example(record, features, name=\"parse_single_example\")\nimg = tf.image.decode_jpeg(parsed_features[\"img_raw\"], channels=3)\nimg.set_shape([180,180,3])\nimg = tf.image.convert_image_dtype(img, tf.float32)\n</code></pre>\n\n<p>then plotting image like that with matplotlib would simply be</p>\n\n<pre><code>with tf.Session() as sess:\n    tmp = sess.run(img)\n    plt.imshow(tmp)\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "222641": "Hi, this is the first time that I deal with this size of dataset . So far my strategy is not successful due to memory and disk limit. I was wondering if someone would like to share how to process the train data?\n\nMy current strategy summarizes as follows:\n\n1. create a VM instances on Google Cloud\n\n2. download train.bson and test.bson data to the vm\n\n3. load bson data into python using Iversion's kernal at https://www.kaggle.com/inversion/processing-bson-files\n\n4. For each image, I was trying to reshape into 1d array and combine it with product_id and category_id. put them as an row in my numpy array.\n\n5. Due to the memory limit, my numpy array cannot hold entire training set. Instead, I was trying to slice the training set into chunks and save each chunk as a separate numpy array and store it into hdf5. However, the corresponding hdf5 can be way much larger than the bson file and even with 50,000 samples, it has been around 40 GB.\n\nCan someone help me with the data management and processing?",
    "222671": "I think dealing with the huge amount of data is as much a challenge as doing the actual classification. ;-)",
    "222697": "vfdev5 spared me some work to the approach I wanted with his kernel: https://www.kaggle.com/vfdev5/random-item-access \n\nDid you ever had to work with data image augmentation? Think about it in the same way, but here you have a single file (much more manageable also, 60Gb) in which you can come pick a fixed number of samples (dependant of your memory) as much as you want (think of it as epochs), acting like a generator.",
    "223853": "Are you using Keras? If so, have you tried fit_generator? It's a way to train in shunks...",
    "224039": "If you dump data into compressed tfrecord files and work with tensorflow you can get down with train dataset to 47GB.",
    "224292": "Hi Marcin, thanks very much for your post. I was trying to use your strategy and almost succeed. However, I keep got error from read the data from tfrecord files. I was trying to fix it for hours but had no luck. My code is as follows. Can you please take a look and help me with where I am wrong? Thanks\n\nThe first block works fine. It reads the bson data and dump it to TFRecords\n\n    def _int64_feature(value):\n      return tf.train.Feature(int64_list=tf.train.Int64List(value=[value]))\n\n    def _bytes_feature(value):\n      return tf.train.Feature(bytes_list=tf.train.BytesList(value=[value]))\n\n    train_filename = 'train.tfrecords'\n    writer = tf.python_io.TFRecordWriter(train_filename)\n    \n    data = bson.decode_file_iter(open('../data/train_example.bson','rb'))\n    with open('../data/train_example_array.txt','ab') as f:\n        for c,d in enumerate(data):\n            product_id = d['_id']\n            category_id = d['category_id']\n            for e,pic in enumerate(d['imgs']): # loop through each picture \n                img = imread(io.BytesIO(pic['picture'])) # output is a 180*180*3=97200 array\n                label = category_id\n                print(img.shape)\n                feature = {'train/label': _int64_feature(label),\n                           'train/image': _bytes_feature(tf.compat.as_bytes(img.tostring()))\n                example = tf.train.Example(features=tf.train.Features(feature=feature)\n                writer.write(example.SerializeToString())\n    writer.close()\n    \nHowever, I get error from the following read from the following reading part. The error says\n\n INFO:tensorflow:Error reported to Coordinator:",
    "224294": "1) Maybe have a look at the data types when you write and read the file. I see some descrepancies in using Int64 and Int32. \n\n2) Also when working with tensorflow it is a good practice to name all transofrmations, so debugger will point you straight to the part that causes error.\n\n3) If you use WITH, you od not need to explicitly close session. \n\nHave a looks at a sample kernel how to read created TFRecord files:\nhttps://www.kaggle.com/mpekalski/reading-tfrecord/",
    "224502": "I finally find a way to get around that error although still confused at where I was wrong. Thanks for sharing your trick.",
    "224624": "Hi Marcin, how you are able to put the entire training data into 47GB compressed TFRecords? I tested my code and found even with 500,000 images it has already taken up 31GB.\n\nmy code is as follows:\n\n    import numpy as np\n    import io\n    import bson\n    import multiprocessing as mp\n    import sys\n    import tensorflow as tf\n    from skimage.data import imread\n   \n    def _int64_feature(value):\n      return tf.train.Feature(int64_list=tf.train.Int64List(value=[value]))\n    def _bytes_feature(value):\n      return tf.train.Feature(bytes_list=tf.train.BytesList(value=[value]))\n    train_filename = '../data/train.tfrecords'\n    opts = tf.python_io.TFRecordOptions(tf.python_io.TFRecordCompressionType.ZLIB)\n    writer = tf.python_io.TFRecordWriter(train_filename,options=opts)\n    data = bson.decode_file_iter(open('../data/train.bson','rb'))\n    for c,d in enumerate(data):\n        product_id = d['_id']\n        category_id = d['category_id']\n        for e,pic in enumerate(d['imgs']): # loop through each picture \n            img = imread(io.BytesIO(pic['picture'])) # output is a 180*180*3=97200 array\n            label = category_id\n            feature = {'label': _int64_feature(label),\n                       'image': _bytes_feature(img.tostring())}\n            # Create an example protocol buffer\n            example = tf.train.Example(features=tf.train.Features(feature=feature))\n            # Serialize to string and write on the file\n            writer.write(example.SerializeToString())\n        if c % 1000 == 0:\n            print(c)\n    writer.close()",
    "224650": "Wenbo, my code is in kernels https://www.kaggle.com/mpekalski/convert-bson-to-tfrecord\n\nFrom what I see you are decoding the image instead of storing it raw as you read it from BSON file.",
    "225280": "Marcin, instead of decoding, I was trying to store the bytes file to TFRecords. However, when I was trying to read from TFRecords, I got dimension error again. someone told me to add \n\n    img_raw = skimage.io.imread(io.StringIO(img_raw).astype(np.uint8).tostring())\n\nbefore write img_raw to disk (img_raw is pic['picture'] which i guess is already bytes file). I don't get the logic why I need to add that line.\n\nMy code and error is at:\nhttps://github.com/wenbo5565/admin/blob/master/Sample%2BRead%2BCompressed%2BTFRecord.ipynb\n\nAppreciate your help!",
    "225680": "Sorry @Wenbo, but I have very limited amount of time I can spend on Kaggle and I cannot spare it at the moment to debug your code.",
    "225717": "Marcin, I already fixed it up. Thanks for all the help so far.",
    "225893": "I did not manage to create any model working with it yet, but in general tensorflow provides means for processing images. After reading stream of bytes you should use **tf.image.decode_jpeg** that returns a **tf.uint8** tensor but we want float32 as input to the model, so we have to convert it.\n\n\n    features =  {\n                 'height': tf.FixedLenFeature((1), tf.int64),\n                 'width': tf.FixedLenFeature((1), tf.int64),\n                 'img_raw': tf.FixedLenFeature((), tf.string),\n                 'category_id': tf.FixedLenFeature((1), tf.int64),\n                 'product_id': tf.FixedLenFeature((1), tf.int64)                 \n                }\n    parsed_features = tf.parse_single_example(record, features, name=\"parse_single_example\")\n    img = tf.image.decode_jpeg(parsed_features[\"img_raw\"], channels=3)\n    img.set_shape([180,180,3])\n    img = tf.image.convert_image_dtype(img, tf.float32)\n\nthen plotting image like that with matplotlib would simply be\n    \n    with tf.Session() as sess:\n        tmp = sess.run(img)\n        plt.imshow(tmp)"
  },
  "source": "meta"
}