{"cells":[{"metadata":{"_uuid":"3a409c468fbc421edd8f6337737f159fbc049d45"},"cell_type":"markdown","source":"The training ata for this competition is provided as a several hundred million line CSV with two columns. I find this a rather awkward data format to deal with. \n\nSimply attempting to load the csv via pandas read_csv takes a long time and isn't very memory efficient. When beginning to experiment with this data in kaggle kernels the pd.read_csv caused a dead kernel when the compute instance ran out of memory. Instead of suffering through a long running custom python loop to do all my feature engineering I have decided to reformat the training data to make the loading faster, more memory efficient and formatted in a way that more closely mimics the way the data is used to generate predictions on the testing segments. \n\nSince the \"time_to_failure\" column varies only very slowly it really needn't be provided for every single row separately, this is doubly true since we have been asked to predict only one such time per stretch of 150,000 acoustic data points in the test segements. By storing just one time_to_failure value per 150,000 training data rows we can save roughly a factor of 2 in memory usage right away, not to mention the extra convenience of having our training and testing data in a similar format. \n\nThe acoustic data has a maximum dynamic range which is small enough to allow us to represent it with arrays of type int16 which saves us a factor of 4 in terms of memory footprint versus storing the data as int64 (this by itself was not enough to. \n\nFinally we can store the data out in the hdf5 data format which will dramatically improve the time required to load the data from disk (roughly 1,000x speedup relative to a pd.read_csv call).\n\nYou can use this alternate format for the training data in your kernels by adding the output of this kernel as an additional data source (see https://www.kaggle.com/product-feedback/45472 )"},{"metadata":{"_uuid":"d91d37fb29d5e82bf08ba59f66f10721ee0bf899","trusted":true},"cell_type":"code","source":"import os\nimport time\nimport h5py\nimport numpy as np\nimport pandas as pd","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3c8807a45487ae4e68a2febb55ae13e7828bbe70"},"cell_type":"markdown","source":"count the number of lines in the training file."},{"metadata":{"_uuid":"6a3edc1eb179173587938bf34cab78883f4288c9","trusted":true},"cell_type":"code","source":"!wc -l ../input/train.csv","execution_count":null,"outputs":[]},{"metadata":{"trusted":false,"_uuid":"fe3b21f2955ffcc301b5751bba03c8fa836b748b"},"cell_type":"code","source":"n_lines_train = 629145481","execution_count":null,"outputs":[]},{"metadata":{"trusted":false,"_uuid":"c7aeffb4b7d1b99ecaa9f93283e0b41a4f5192d6"},"cell_type":"code","source":"chunk_size = 150000\nn_segments = (n_lines_train-1)//chunk_size\nn_segments","execution_count":null,"outputs":[]},{"metadata":{"trusted":false,"_uuid":"7c498bae67e238763dbcbb1b419f957bd6ca5f2b"},"cell_type":"code","source":"leftover = n_lines_train - n_segments*chunk_size\nleftover","execution_count":null,"outputs":[]},{"metadata":{"trusted":false,"_uuid":"383dc6cc55077f2711d26c662052bf5938888b2c"},"cell_type":"code","source":"leftover/n_lines_train","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ffe17997b75f50c3448011abaeeb594d252606a1"},"cell_type":"markdown","source":"Unfortunately the training data doesn't neatly divide into an integer number of the test segment size chunks. But when dealing with 600+ million time points a few tens of thousand more or less probably won't make much of a difference (it makes up just one ten thousandth part of the training data). So I will just ignore the last few training data time points. "},{"metadata":{"trusted":false,"_uuid":"0699fc9885f6d30b9071d0f07082caef15a6ff98"},"cell_type":"code","source":"input_dir = \"../input\"\noutput_dir = \"\"","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"71fcc65d0d40ecd8db3cf4e2238ae7cc1d16707d"},"cell_type":"markdown","source":"we create the hdf5 file and then will create named datasets within the file and then iterate over the csv and fill in the dataset rows one by one. We also could have first loaded the data into numpy arrays and then simply assigned the arrays directly out to the hdf5 file which is often a more convenient interface, but doing it this way allows us to deal with files which are much too large to fit directly into memory."},{"metadata":{"trusted":false,"_uuid":"b5ec2aa8c7bdc4b12f82ddcb7df160c544fa68c7"},"cell_type":"code","source":"#create the hdf5 file\nh5_file = h5py.File(os.path.join(output_dir, \"train.h5\"), \"w\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":false,"_uuid":"c3086234512906f73cf2717124470518ff5ad7c9"},"cell_type":"code","source":"#create datasets within the top level group in that file\nsound_dset = h5_file.create_dataset(\"sound\", shape=(n_segments, chunk_size), dtype=np.int16)\nttf_dset = h5_file.create_dataset(\"ttf\", shape=(n_segments,), dtype=np.float32)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"61967145c3d546b79793589a7e0c67018d500777","trusted":false},"cell_type":"code","source":"#iterate over all 629 million lines and save them out in chunks of 150,000\n#this takes a while ...\nchunk_size = 150000\nlines_to_read = chunk_size*n_segments\n\nprinted_warning = False\nwith open(os.path.join(input_dir, \"train.csv\")) as f:\n    x_stack, y_stack = [], []\n    last_ttf = np.inf\n    hdr = f.readline()\n    for line_idx in range(lines_to_read):\n        cx, cy = f.readline().split(\",\")\n        cx = int(cx)\n        if np.abs(cx) > 32767:\n            if not printed_warning:\n                printed_warning = True\n                print(\"line {} is too big to be an int16\".format(line_idx+1))\n        cy = float(cy)\n        if cy < last_ttf:\n            last_ttf = cy\n            y_stack.append(cy)\n        x_stack.append(cx)\n        if line_idx % chunk_size == chunk_size-1:\n            sound_dset[line_idx//chunk_size] = np.array(x_stack).astype(np.int16)\n            ttf_dset[line_idx//chunk_size] = np.mean(y_stack)\n            x_stack, y_stack = [], []\n            last_ttf = np.inf","execution_count":null,"outputs":[]},{"metadata":{"trusted":false,"_uuid":"960305cb77c68ace0b12a75de59dc0022b5ca0fd"},"cell_type":"code","source":"h5_file.close()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7ae923dd383229359338c0ce218184ab94f260de"},"cell_type":"markdown","source":"Now that we have a nice hdf5 file lets do some comparisons versus the csv file. First off the h5 file on disk takes up only 1.2 Gb versus 8.9 Gb for the uncompressed csv. Most of this difference can be attributed to the fact that the time_to_failure column takes a lot of characters to express as a character string and we are storing just 1/150,000th as many numbers and in a binary format. "},{"metadata":{"_uuid":"87a65f014152c062db8178c076f87afb8c8a132d"},"cell_type":"markdown","source":"What about the time to read in the data from disk? Lets load the first 100 segments from our hdf5 file and compare that to using pandas.read_csv on the raw data."},{"metadata":{"trusted":false,"_uuid":"72fa351ef796c4da196dc8d91a8232adad0ce3f7"},"cell_type":"code","source":"start_time = time.time()\nhf = h5py.File(os.path.join(output_dir, \"train.h5\"))\nsegs = np.array(hf[\"sound\"][:100])\nseg_ttf = np.array(hf[\"ttf\"][:100])\nhf.close()\nend_time = time.time()\nprint(\"{} seconds\".format(end_time-start_time))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4b51f855f9a9a0068e26656f7252dc77e1a16cf4"},"cell_type":"markdown","source":"now compare this with the time to read in the data from the csv using pandas."},{"metadata":{"trusted":false,"_uuid":"06ce50c03d3ccc0f735946751f9c06cb6fd22dad"},"cell_type":"code","source":"start_time = time.time()\ndf = pd.read_csv(\n    os.path.join(input_dir, \"train.csv\"), \n    nrows=chunk_size*100,#limit to the first 100 segments \n    dtype={\"acoustic_data\":np.int16, \"time_to_failure\":np.float32}\n)\nend_time = time.time()\nprint(\"{} seconds\".format(end_time-start_time))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ff7b1637d1db29169aa1d42b59049949a035c491"},"cell_type":"markdown","source":"After putting the data into the hdf5 file the read time is fast, the memory usage is reduced and the training data is neatly formatted into 150,000 length segments and single time labels just like at testing time. \n\nTo maximize convenience we can also turn the test segment data into a hdf5 file with the same format. "},{"metadata":{"trusted":false,"_uuid":"d04c76d887c5b423a4ac8ba282b32a207fd86526"},"cell_type":"code","source":"test_files = os.listdir(os.path.join(input_dir, \"test\"))\n\nwith h5py.File(os.path.join(output_dir, \"test.h5\"), \"w\") as h5_file:\n    sound_dset = h5_file.create_dataset(\"sound\", (len(test_files), chunk_size))\n    seg_ids = []\n\n    for fidx, fname in enumerate(test_files):\n        cdata = pd.read_csv(os.path.join(input_dir, \"test\", fname), dtype=np.int16)[\"acoustic_data\"].values\n        sound_dset[fidx] = cdata\n        seg_ids.append(fname.split(\".\")[0])\n\n    h5_file[\"seg_id\"] = np.array(seg_ids).astype(np.string_)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"844f885a804b1e3412170c328ad0c11f2ec27fc7"},"cell_type":"markdown","source":"While HDF5 is fantastic for storing numerical data it can be somewhat painful to use for storing text data. Unfortunately HDF5 likes storing string arrays as ascii and numpy likes string arrays to be either of object type, fixed length ascii (which works for hdf5 just fine) or unicode (which doesn't work with hdf5). \n\nSo when we write out the seg_id's we need to use np.string_ data format which will encode the strings as ascii. Then when we load the seg_id's back in we need to explicitly cast the numpy array back to the \"str\" type so that python will treat the resulting array as strings instead of as type \"bytes\". "},{"metadata":{"trusted":true,"_uuid":"038303944d3d059a311cef8419dcdba25e055527"},"cell_type":"code","source":"#loading back in the test data segment strings\nhf = h5py.File(os.path.join(output_dir, \"test.h5\"))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6545086f4c18e121ee68b7df5f02a4cad843d9e3"},"cell_type":"code","source":"#without a .astype(str) the resulting array is of type bytes\nseg_ids = np.array(hf[\"seg_id\"])\nseg_ids","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d9eccecdfcbe94d3983f8b1137be4367549e6f5a"},"cell_type":"code","source":"#adding a astype call gets us nice unicode strings\n#that play well with python 3\nseg_ids = np.array(hf[\"seg_id\"]).astype(str)\nseg_ids","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c9f8836c11b1a9aae19e4662e6968a508d8b6a93"},"cell_type":"code","source":"hf.close()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"28777ee8cc4948889a079a4df18fbc9f0fe9eb12"},"cell_type":"markdown","source":"Happy kaggling."}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.6.3"}},"nbformat":4,"nbformat_minor":1}