{"cells":[{"metadata":{"_uuid":"5f83661a006a87f4a92c775984e03ea3273df395"},"cell_type":"markdown","source":"Hello everyone, so this competition I couldn't get the idea I wanted working but I came to some interesting conclusions and felt they were worth sharing. I hope for this kernel to be a small antithesis to the last second blending kernels that screw up the rankings without offering anything to the community. One aspect I really spent a great deal of time on was the image data. I am sure the top teams have found some great way to handle the large quantity of images and an inteligent model or features to extract the information from this, but I figured I would share the research process I have gone through and possibly help someone along so they dont have to do some of the fundamental testing I went through.  "},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"collapsed":true},"cell_type":"markdown","source":"First off lets look at the process of loading in 1.3 million images. We have many options here. "},{"metadata":{"_uuid":"ea84de0a4fa1017dbacdfbffcfea868fb64d9692"},"cell_type":"markdown","source":"Speed on this is very important. You don't want to have idle gpu and be waiting on disk to load in the images. I actually never tried with the zip file. I just went for the unzipped files and only ever dealt with those. That may be the way to go, I do not know. My assumption was that repeatedly unzipping the files would be a waste. Even if that the case, it is very likely someone will come across a problem where they don't received the images zipped like that and this will be relevant to them still. Anyway, lets start by loading in just the images and seeing what various libraries can do in terms of loading in the images"},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"65f4e6cb5704d2b645a30b40387379a3578a896b"},"cell_type":"code","source":"#I don't think this line needs explaining\nimport pandas as pd","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c73f8b9ca19d0bd22f316e8fa72f65b1c59d0ba2","collapsed":true},"cell_type":"code","source":"#load the train csv and grab the images\nimages = pd.read_csv(\"../input/train.csv\")[[\"image\"]]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"e3d4e99c7085d4de3c21d3ec21b8402713877928"},"cell_type":"code","source":"#let's measure some time\nimport time","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"46441f7a3044341c0b1c4a7785f8d3c488cf2aac"},"cell_type":"markdown","source":"We'll start with PIL to load the images. I am sure many of you are familiar with this library. "},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"9d15db421df78fd12dcc84d8442a838cc908e68c"},"cell_type":"code","source":"#I'm sure importing open is probably a terrible pythonic sin, but yolo\nfrom PIL.Image import open as PILRead","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f5996225ebe3e54788882f63fa8f0e01d3900236"},"cell_type":"markdown","source":"Sorry but this code won't work on a kernel as it is all under the assumption that you have already unzipped the jpg folders. \n\nFirst thing we will do is measure time to iterate through the first thousand images and see how long it takes simply to load the image."},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"e58aee55cf4038c6453c2c0a68740ba570837b6d","_kg_hide-output":true},"cell_type":"code","source":"start = time.time()\nfor image in images[0:1000]:\n        img = PILRead('./data/competition_files/train_jpg/' + str(image) + \".jpg\")\nprint(time.time() - start)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3cbcb13ae2dba22e258921a902b6531e5a8b01e1"},"cell_type":"markdown","source":"    ---------------------------------------------------------------------------\n    FileNotFoundError                         Traceback (most recent call last)\n    <ipython-input-41-33e0c0720396> in <module>()\n          1 start = time.time()\n          2 for image in X[\"image\"][0:1000]:\n    ----> 3         img = PILRead('./data/competition_files/train_jpg/' + str(image) + \".jpg\")\n          4 print(time.time() - start, i/1000)\n\n    c:\\users\\magic\\appdata\\local\\programs\\python\\python35\\lib\\site-packages\\PIL\\Image.py in open(fp, mode)\n       2546 \n       2547     if filename:\n    -> 2548         fp = builtins.open(filename, \"rb\")\n       2549         exclusive_fp = True\n       2550 \n\n    FileNotFoundError: [Errno 2] No such file or directory: './data/competition_files/train_jpg/nan.jpg'\n"},{"metadata":{"_uuid":"81be36252b072bc7dc5a4a2d6a6261dec6eab570"},"cell_type":"markdown","source":"Drats. Errors. Looks like we have some missing images. We'll just return an array of zeros of our target size if we cant load an image properly I guess. Let's also keep track of how many aren't able to be loaded for whatever reason.\nSecond Try:"},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"3b76a8844ecd9f2d56c32ddc2782ceb950ebf889","_kg_hide-output":true},"cell_type":"code","source":"i = 0\nstart = time.time()\nfor image in X[\"image\"][0:1000]:\n    try:\n        img = PILRead('./data/competition_files/train_jpg/' + str(image) + \".jpg\")\n        i += 1\n    except:\n        img = np.zeros(shape = (224, 224, 3))\nprint(time.time() - start, i/1000)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c830a4cde044cd77d478e0ca7e975c4f47e477e2"},"cell_type":"raw","source":"0.19799304008483887 0.921\nLooks like my favorite tool to make the bad men go away, the try clause, saves the day again. Hey that's pretty fast. Missing that many images isn't perfect but that's just part of the data science gig. PIL looks like a decent candidate"},{"metadata":{"_kg_hide-output":true,"trusted":true,"collapsed":true,"_uuid":"6446e61bb8aed1d3b7f5e0b424f7d379a51e6bab"},"cell_type":"code","source":"from cv2 import imread as cvimread","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"9da9b2af1f5a08ae71a634f882de1bc0c587816e","_kg_hide-output":true},"cell_type":"code","source":"start = time.time()\nfor image in X[\"image\"][0:1000]:\n    img = cvimread('./data/competition_files/train_jpg/' + str(image) + \".jpg\")\nprint(time.time() - start)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"606aef2acc9b4eb73bdfceb6d1164025d203dd3c"},"cell_type":"markdown","source":"2.5197324752807617. That's way slower, why is that? and why didn't it error out like the last one? Time for exploration. Lets check some shapes"},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"c4ad1eab8c9fb7ef7c85011bb84495b4153c760d","_kg_hide-output":true},"cell_type":"code","source":"start = time.time()\nfor image in X[\"image\"][0:1000]:\n    img = cvimread('./data/competition_files/train_jpg/' + str(image) + \".jpg\")\n    print(img.shape)\nprint(time.time() - start)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"fe0e2475dc02a10000085271d1133484cc01a07b"},"cell_type":"markdown","source":"    (480, 358, 3)\n    (480, 360, 3)\n    (360, 392, 3)\n    (360, 360, 3)\n    (360, 640, 3)\n    (480, 360, 3)\n    (360, 480, 3)\n    (480, 360, 3)\n    (480, 360, 3)\n    (480, 270, 3)\n    (480, 270, 3)\n    (480, 360, 3)\n    (463, 360, 3)\n    (360, 480, 3)\n    (360, 480, 3)\n    (480, 320, 3)\n    (360, 640, 3)\n    (480, 360, 3)\n    (480, 270, 3)\n    ---------------------------------------------------------------------------\n    AttributeError                            Traceback (most recent call last)\n    <ipython-input-48-7289f6063999> in <module>()\n          2 for image in X[\"image\"][0:1000]:\n          3     img = cvimread('./data/competition_files/train_jpg/' + str(image) + \".jpg\")\n    ----> 4     print(img.shape)\n          5 print(time.time() - start)\n\n    AttributeError: 'NoneType' object has no attribute 'shape'\n    \nLooks like cv2 is returning None instead of throwing an error. Still doesn't explain why it is so slow. Let's make sure it is returning what we want with both cv2 and PIL by checking the type"},{"metadata":{"_kg_hide-output":true,"trusted":true,"_uuid":"6c8ef7a7a8a99cd2916d821cd74f71a060a681b9","collapsed":true},"cell_type":"code","source":"type(img)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b396b9511876221fda9d17e5884d6d95e484e05e"},"cell_type":"markdown","source":"    numpy.ndarray\nPerfect. cv2 seems to be returning what we expect. An array with all of the pixel values. Let's be lazy and copy and paste the code from above for PIL and check again. "},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"74920cb18ab7625d2c823dc5dee3672638351b77","_kg_hide-output":true},"cell_type":"code","source":"i = 0\nstart = time.time()\nfor image in X[\"image\"][0:1000]:\n    try:\n        img = PILRead('./data/competition_files/train_jpg/' + str(image) + \".jpg\")\n        i += 1\n    except:\n        img = np.zeros(shape = (224, 224, 3))\nprint(time.time() - start, i/1000)\ntype(img)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c47cce7f104823dcb538a1baa64b8d6f4451990a"},"cell_type":"markdown","source":"    PIL.JpegImagePlugin.JpegImageFile\n\nHuh, you're not my mom? On further research it looks like PIL isn't really doing what I intended, but we can try to convert this object into what we need, an array, using a keras utility.\n"},{"metadata":{"trusted":true,"_uuid":"d1aad6373d2e149e539126dd8cdb1fb96edd0860","_kg_hide-output":true,"collapsed":true},"cell_type":"code","source":"from keras.preprocessing.image import img_to_array","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"e15484aca5aaeba0135800ec224fbf1e662b2e1b","_kg_hide-output":true},"cell_type":"code","source":"i = 0\nstart = time.time()\nfor image in X[\"image\"][0:1000]:\n    try:\n        img = img_to_array(PILRead('./data/competition_files/train_jpg/' + str(image) + \".jpg\"))\n        i += 1\n    except:\n        img = np.zeros(shape = (224, 224, 3))\nprint(time.time() - start, i/1000)\ntype(img)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"228eddb95cfbf15f06df386ac4df6b6716fbb6e8"},"cell_type":"markdown","source":"    3.4392542839050293 0.921\n    numpy.ndarray\nThere we go. More what we wanted but looks like it is actually slower than cv2 now. There might exist a faster image to array tool, but I didnt do any additional research on this. Let's look and compare with some other image loaders though"},{"metadata":{"_uuid":"7ba145153db1af8ab013f31567a967138f193874"},"cell_type":"markdown","source":"Trusty scipy has some tool that looks like it might work"},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"5d3c3d120d839b2d71447ecdeba889ff5895b17b","_kg_hide-output":true},"cell_type":"code","source":"from scipy.misc import imread as scimread\ni = 0\nstart = time.time()\nfor image in X[\"image\"][0:1000]:\n    try:\n        img = (scimread('./data/competition_files/train_jpg/' + str(image) + \".jpg\"))\n        i += 1\n    except:\n        img = np.zeros(shape = (224, 224, 3))\nprint(time.time() - start, i/1000)\ntype(img)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"de74d8678187298dc1a2a2122b83bebd7e7e3c68"},"cell_type":"markdown","source":"    2.7287728786468506 0.921\nPretty good but gives me a deprecation warning though. Let's try what they recommend."},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"b784da6c7df7e4c91b7fbca73d5d2703a7299a53","_kg_hide-output":true},"cell_type":"code","source":"from imageio import imread as ioimread\n\ni = 0\nstart = time.time()\nfor image in X[\"image\"][0:1000]:\n    try:\n        img = ioimread('./data/competition_files/train_jpg/' + str(image) + \".jpg\")\n        i += 1\n    except:\n        img = np.zeros(shape = (224, 224, 3))\nprint(time.time() - start, i/1000)\ntype(img)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0bc99fa2e08b11fe2e6ce5dacb4e5014b4edef71"},"cell_type":"markdown","source":"    2.8476409912109375 0.921\n    imageio.core.util.Image\nDid some research and seems like this bizarre type would work just fine like a numpy array, but weird type and slightly slower. Probably not worth."},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"81914f025a2a001f9beb8b4849d7baaf7f752f13","_kg_hide-output":true},"cell_type":"code","source":"from matplotlib.image import imread as matimread\ni = 0\nstart = time.time()\nfor image in X[\"image\"][0:1000]:\n    try:\n        img = (matimread('./data/competition_files/train_jpg/' + str(image) + \".jpg\"))\n        i += 1\n    except:\n        img = np.zeros(shape = (224, 224, 3))\nprint(time.time() - start, i/1000)\ntype(img)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"19718fab664fe879e80b833c282d712d399da5ae"},"cell_type":"markdown","source":"matplotlib's library looks like it does pretty well and returns what we expect but still marginally slower than cv2. We are dealing with relatively small margins but still. If we are pushing this to so many images it might be worth it to care about these small margins. \n2.780322313308716 0.921\nnumpy.ndarray"},{"metadata":{"_uuid":"5d6b474269e8c86b48308821ba907b482e4e8e4c"},"cell_type":"markdown","source":"Now we'll try keras's tools. Those must be fast. They've surely considered and curated the fastest implementations already. We'll chain together their load_img and img_to_array to get an array like we expect"},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"31008b75ae81423929e2e12b872fe6dceb154265","_kg_hide-output":true},"cell_type":"code","source":"from keras.preprocessing.image import load_img\ni = 0\nstart = time.time()\nfor image in X[\"image\"][0:1000]:\n    try:\n        img = img_to_array(load_img('./data/competition_files/train_jpg/' + str(image) + \".jpg\"))\n        i += 1\n    except:\n        img = np.zeros(shape = (224, 224, 3))\nprint(time.time() - start, i/1000)\ntype(img)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f118255d4ffbebab07efaf5bd65501368a5f1c60"},"cell_type":"markdown","source":"    3.6963675022125244 0.921\n    numpy.ndarray\nWell looks like cv2 it is. All we have to do is deal with the small quirk where it returns None sometimes. "},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"b00da74fbe32e337da7792b4a698c12a6577841e","_kg_hide-output":true},"cell_type":"code","source":"start = time.time()\nfor image in X[\"image\"][0:1000]:\n    img = cvimread('./data/competition_files/train_jpg/' + str(image) + \".jpg\")\n    try:\n        img.shape\n    except:\n        img = np.zeros(shape = (224, 224, 3))\nprint(time.time() - start)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a8e83ede64ffdcb425df1ce9365d8e301ebea639"},"cell_type":"markdown","source":"    2.6547305583953857\nNot the most eloquent solution, but it's the fastest I could come up with. Started reading a book to up my python skills, mentions that throwing expection is preferable to returning none. This is a prime example. "},{"metadata":{"_uuid":"0f57973f24595a0a798d50174f592f32800ebd5a"},"cell_type":"markdown","source":"Now lets rewind back to something that we have glossed over. If we look at the images they arent' the same size. The resolution is variable. I don't know of a way to deal with images of variable size in a NN. It may exist, but it'd be easiest to resize them to all be the same size. Let's do 224, 224, 3 as that seems to be what some of the major pretrained CNN's come as. "},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"e0e686d4a0f261ec8b9eeda24af73614327c8766","_kg_hide-output":true},"cell_type":"code","source":"from cv2 import resize\nstart = time.time()\nfor image in X[\"image\"][0:1000]:\n    img = cvimread('./data/competition_files/train_jpg/' + str(image) + \".jpg\")\n    try:\n        img = resize(img, dsize=(224, 224), interpolation=cv2.INTER_CUBIC)\n        img.shape\n    except:\n        img = np.zeros(shape = (224, 224, 3))\nprint(time.time() - start)\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"29618a97ca516ce6fe97ffd446000066d01e6b66"},"cell_type":"markdown","source":"    3.6379148960113525"},{"metadata":{"_uuid":"43d3dadbaaaac4e7d8d39104ce2d94726139e659"},"cell_type":"markdown","source":"Using the same previous code snippet plus the resize function from cv2. Shows that this resize process adds some time. This is a very simple way to rescale the image. I apologize for whoever wrote the kernel because I can't find it anymore but I also utilized someones code to try to pad the sides so all dimensions were the same before reducing the image to 224,224 because that causes some rectangular images to become squished. Wasn't able to test the hypothesis, but I don't think one is necessarily better than the other because padding the sides also destroys quite a bit of resolution despite keeping the same aspect ratio. "},{"metadata":{"_uuid":"08de4395a467f63bbb3e57619ad2ab6a4e554cea"},"cell_type":"markdown","source":"Anyway, is that the fastest we can do? Well no. Everyone knows for loops are slow. What if we could deal with the data in parallel instead of linearly while maintaining order? Too good to be true? Enter concurrent.futures or any other number of multithreading tools. For this one I will focus on concurrent.futures because the map function is very slick and easy. "},{"metadata":{"_uuid":"9fe9bb13fdea6a4c9274fbdd0b95c44442fb972a"},"cell_type":"markdown","source":"Let's rewrite our previous snippet as a function so we can use map. "},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"d79d614c409918260b890f32a8525a40df277766","_kg_hide-output":true},"cell_type":"code","source":"import concurrent.futures\ndef resize_img(image):\n    img = cvimread('./data/competition_files/train_jpg/' + str(image) + \".jpg\")\n    try:\n        img = resize(img, dsize=(224, 224), interpolation=cv2.INTER_CUBIC)\n        img.shape\n    except:\n        img = np.zeros(shape = (224, 224, 3))\n    return img","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"314bec4b673388b63c644c41c20018273b2b55de","_kg_hide-output":true},"cell_type":"code","source":"start = time.time()\nresized_imgs = []\nwith concurrent.futures.ThreadPoolExecutor(max_workers = 16) as executor:\n    for value in executor.map(resize_img, X[\"image\"][0:1000]):\n        resized_imgs.append(value)\nprint(time.time() - start)      ","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"53b9233796597e0c611d80b4b558a6a0a75abf2e"},"cell_type":"markdown","source":"   0.49677205085754395"},{"metadata":{"_uuid":"604279e5ca081529012f2e5b71d5f4d907952def"},"cell_type":"markdown","source":"concurrent.futures map is very very cool and easy to use. It allows us to take a list like we have and then gobble it up in parallel with our function and then it does the work of piecing it together in the correct order. max_workers is a parameter you can toy with based on your system configuration. Things to consider when changing that: amount of cores and available memory. If you queue too many things up then you might throw too much in your memory and might spill over into disk which will greatly slow things down or it may be more cores than you have available. I show the images being appended to a list. This likely isn't what you want to do, but I'm too lazy to make a better example. \n\nIf you want to dig into this a little more there is a very good stack overflow post looking at some of the performance considerations and how you might use submit instead of map sometimes. https://stackoverflow.com/questions/42074501/python-concurrent-futures-processpoolexecutor-performance-of-submit-vs-map\nAlso, just to note if you are on Windows it is technically possible to use the ProcessPoolExecutor instead of ThreadPoolExecutor, but I'm going to suggest you spend the time just installing Linux if you are really wanting to go down that path. This is actually something Keras has decided to just totally avoid. https://github.com/keras-team/keras/pull/8662\n\nIf you are on Linux actually a lot of the problems we have been focusing on go away because you can use the multiprocessing option in Keras with a generator and number of threads and queue size all built in, but I thought it would be valuable to share anyway. So we're done, right? But what happens if we resize the images before loading them in and write them to a separate folder. How fast would that be? "},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"ac6b6f5b626536fcb12430c9073ef4e328765c16","_kg_hide-output":true},"cell_type":"code","source":"import concurrent.futures\ndef resize_img(image):\n    img = cvimread('./data/competition_files/train_jpg_resized/' + str(image) + \".jpg\")\n    try:\n        img.shape\n    except:\n        img = np.zeros(shape = (224, 224, 3))\n    return img","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"5dfea6c9eaab77ac128e8675fc681b5f3efb911f","_kg_hide-output":true,"collapsed":true},"cell_type":"code","source":"start = time.time()\nresized_imgs = []\nwith concurrent.futures.ThreadPoolExecutor(max_workers = 16) as executor:\n    for value in executor.map(resize_img, X[\"image\"][0:1000]):\n        resized_imgs.append(value)\nprint(time.time() - start)    ","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d9034ce1e6c6329379dc2fe908dde10b6073dc8a"},"cell_type":"markdown","source":"    0.16491961479187012\nEven better.  So if we just resized the images beforehand and put them in a new folder we can load in the files stupid quick. We took a process that could take 3+ seconds per round if we tried to load in serially and resize on the fly down to .165 seconds. This is crucial for quick prototyping and testing of new ideas. \n\nThis is what I ended up using for most of my attempts at a model that utilized the images directly, but I also went down the rabbit hole of rewriting the images in HDF5 for faster reading. I will post that research tomorrow when it isn't 2:30am for me. Also have some work I did regarding Keras's flow_from_directory, but ultimately didn't find a good way to use that. For now I will just focus on sharing the HDF5 stuff because I think that will actually be highly relevant not just to this challenge but to many peoples and the big data they may be dealing with. "}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.5","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}