{"cells":[{"metadata":{"_uuid":"c715d307f6513315cb259b52aa89178a2be9602a"},"cell_type":"markdown","source":"Let's try to understand how data is **distributed over time** in [LANL Earthquake Prediction](https://www.kaggle.com/c/LANL-Earthquake-Prediction) training dataset."},{"metadata":{"_uuid":"bf5e25fd074096a9dc5d1179b882d1373aabca30"},"cell_type":"markdown","source":"Dataset is roughly 9GB. So what would we use to load it?\n\nnumpy.loadtxt is known to be awkwardly slow. On the other hand pandas.read_csv is really clumsy in terms of memory consumption.\nI've tried few additional flags, but still can't guarantee to run this code multiple times."},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import pandas\nimport numpy as np\n\ndt = { 'acoustic_data': 'i2', 'time_to_failure': 'f8' }\ndata = pandas.read_csv(\"../input/train.csv\", dtype=dt, engine='c', low_memory=True)\n\nN = data.shape[0]\nprint(\"Data size\", N)\n#print(\"Data stats\", data.describe())","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d0801070bf5bbdedaddee3229f96fbea8eeb947a"},"cell_type":"markdown","source":"Meantime, test data chunks all have same size of 150k records and doesn't have any time markers. So unless further clarified by competition organizers, we'll have to make some assumptions about time function for test chunks and training data, and their correllation.\n\nAs stated by organizers in [Data description](https://www.kaggle.com/c/LANL-Earthquake-Prediction/data) and [Additional info](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77526):\n> The training data is a single, continuous segment of experimental data.\n\n> The input is a chunk of 0.0375 seconds of seismic data (ordered in time), which is recorded at 4MHz, hence 150'000 data points, and the output is time remaining until the following lab earthquake, in seconds.\n\n> Both the training and the testing set come from the same experiment.\n\nDoes it mean, that training data also recorded at 4MHz? Probably.\nAnd that it also consists of 0.0375 second (150'000 data points) chunks? I don't think so. In this kernel, I'll try to show why.\n\nLet's figure out what does \"continuous segment of experimental data\" mean in reality.\n\nI would read it as: **Time interval (and thus difference between \"time_to_failure\" values) between adjacent data frames is constant**, unless \"time_to_failure\" values point to 2 different earthquakes.\n\nLet's take a closer look into time_to_failure sequence, and validate this interpetation."},{"metadata":{"trusted":true,"_uuid":"45e015c7b3384008bda0b62e5a9ab91be243ba1f","_kg_hide-input":true},"cell_type":"code","source":"digits = 10\n\noldValue = data.iloc[0,1]\nnewValue = data.iloc[1,1]\nframe_diff = round(oldValue - newValue, digits)\noldDiff = frame_diff\nprint(\"Time to failure\", round(oldValue, digits))\nfor i in range(1, 20000):\n    newValue = data.iloc[i,1]\n    newDiff = round(oldValue - newValue, digits)\n    if oldDiff != newDiff:\n        print(\"Time difference changed from\", oldDiff, \"to\", newDiff, \"on frame\", i)\n        oldDiff = newDiff\n    oldValue = newValue\nprint(\"Time to failure\", round(newValue, digits))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ccd9d5dbc8c27c9b483fda82cecabc2ce0301c1d"},"cell_type":"markdown","source":"Code above tests if frame rate is constant or changes over time. As we can see there are gaps before every 4096-th frame.\n\nAnd to be clear on durations of these gaps:"},{"metadata":{"trusted":true,"_uuid":"cb6ddcb66ac2d171a448cb2917ba92af64f13c6e"},"cell_type":"code","source":"chunk_diffs = [\n    round(data.iloc[0,1] - data.iloc[4096,1], digits),\n    round(data.iloc[4095,1] - data.iloc[8191,1], digits)\n]\nprint(chunk_diffs)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ef0c92ef218802c501f42413d9fd16bc1b25aba7"},"cell_type":"markdown","source":"Obviously first continuous data chunk has length of 4095 frames, while next few chunks have length of 4096 frames.\n\nSo, does all dataset consist of chunks of nearly 4096 (2^12) length?\n\nLets make few more calculations."},{"metadata":{"trusted":true,"_uuid":"03148aff3139dc1a0b0455b6d259ee0e676168b4"},"cell_type":"code","source":"M = 4096\nC = N//M+1\nR = C*M - N\nprint(\"Chunk size\", M)\nprint(\"Number of chunks\", C)\nprint(\"Number of incomplete chunks\", R)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"dfa921bb20d0ec0e7e7249dc491df832a6e9933a"},"cell_type":"markdown","source":"Looks quite promising.\n\nSo if our theory is true:\n* There are 153600 chunks of sequential data.\n* Chunks have length of **4096 frames**, with 120 missing frames spreaded among all chunks.\n* Time between frames inside chunk is **1.1e-9**.\n* Time between corresponding (first or last) frames of adjucent chunks is either **0.001 or 0.0011**\n\nLet's validate. Code below **will stop execution if above constrains are not satisfied**. There are, however, few notes to consider:\n* Some \"time_to_failure\" values in source data are rounded to 1e-9 precision losing last digit of frame duration (which is 1.1e-9)\n* If one of the adjucent chunks has less then 4096 frames, and misses few starting or ending frames, then time beetween first frames might be not exactly equal to either 0.001 or 0.0011.\n\nAdditionally, acoustic data and chunks metadata related to **separate earthquakes saved to individual output files**. This would be valuable outcome for me and anybody willing to perform analysis earthquake-wise and **significanlty reduce size** of source data to load.\n\nNote: *Chunk metadata being saved consists of first frame index for each chunk, and time difference between first frame and first frame of previous chunk. First chunk stores initial \"time_to_failure\" instead. This should be enough to restore \"time_to_failure\" values for every frame.*"},{"metadata":{"trusted":true,"_uuid":"50e21074fa4eaa827b02928d478763d02f57fa34","_kg_hide-input":true},"cell_type":"code","source":"chunk_digits = 8\n\nmax_error = 1e-9\nsmall_error_count = 10\nfailure_starts = [[0,0]]\nfailure_index = 0\nchunks = np.zeros((C, 2))\nchunk_index = 0\nstartValue = data.iloc[0,1]\nchunks[chunk_index] = [0, round(startValue, digits)]\nprint(\"Time to\", failure_index, \"earthquake\", round(startValue, digits))\nfor i in range(1, N):\n    j = i - chunks[chunk_index,0]\n    newValue = data.iloc[i,1]\n    newDiff = round(startValue - newValue, digits)\n    frame_error = newDiff - round(frame_diff * j, digits)\n    if newDiff < 0:\n        print(\"New earthquake after\", chunk_index+1, \"chunks\", i-failure_starts[failure_index][0], \"samples\")\n        data.iloc[failure_starts[failure_index][0]:i,0].to_csv('failure{0}_data.zip'.format(failure_index), index=False, compression=\"zip\")\n        np.savetxt('failure{0}_chunks.csv'.format(failure_index), chunks[failure_starts[failure_index][1]:chunk_index+1], fmt='%d, %f', header='Frame, Time', comments='')\n        print(\"Time to\", failure_index, \"earthquake\", round(startValue, digits))\n        failure_index += 1\n        chunk_index += 1\n        failure_starts.append([i, chunk_index])\n        startValue = newValue\n        chunks[chunk_index] = [i, round(startValue, digits)]\n        print(\"Time to\", failure_index, \"earthquake\", round(startValue, digits))\n    elif round(newDiff, chunk_digits) in chunk_diffs:\n        # \n        chunk_index += 1\n        chunks[chunk_index] = [i, newDiff]\n        startValue = newValue\n        if chunk_index % 1000 == 0:\n            print(\"Chunk\", chunk_index)\n    elif frame_error != 0:\n        if small_error_count > 0:\n            print(\"Unexpected frame\", failure_index, chunk_index, j, \"duration\", newDiff, \"expected\", round(frame_diff * j, digits))\n            small_error_count -= 1\n        if frame_error > max_error:\n            print(\"Prediction error, stopping execution\")\n            break\n\ndata.iloc[failure_starts[failure_index][0]:,0].to_csv('failure{0}_data.zip'.format(failure_index), index=False, compression=\"zip\")\nnp.savetxt('failure{0}_chunks.csv'.format(failure_index), chunks[failure_starts[failure_index][1]:], fmt='%d, %f', header='Frame, Time', comments='')\nprint(\"Time to\", failure_index, \"earthquake\", round(newValue, digits))","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"markdown","source":"Missing starting frames for some chunks are completely fine, and shouldn't have noticable impact on predictions quality.\n\nTime differences between chunks seem to follow nearly cyclic pattern with cycle consisting of 25 chunks. But it sometimes changes throughout training data, so additional investigation might be necessary to make any assumptions about nature of time gaps between chunks.\n\nNone of the gaps is enough to put 0.0375 second (or 150'000 frames) inside. So if training data somehow overlaps with tests data, there will be overlapping series of accoustic data values at least of (nearly) 4096 size.\n\nPlease, note, that if **test data represents long continuous time segments** as stated, while **training data consists only of short continuous chunks**, with seemingly **different frame rate** inside each chunk, it might **significantly affect prediction quality** for solutions trained only on training dataset."},{"metadata":{"trusted":true,"_uuid":"741b147cc9c67e930c5f35ba6f78e5fe2854976d"},"cell_type":"code","source":"failure_count = failure_index + 1\nchunk_count = chunk_index + 1\nprint(\"Chunk count expected\", C, \"actual\", chunk_count)\nfor fi in range(failure_count):\n    first_frame = failure_starts[fi][0]\n    last_frame = int(chunks[(failure_starts[fi+1][1] if fi < failure_count - 1 else chunk_count) - 1, 0])\n    first_ttf = data.iloc[first_frame,1]\n    last_ttf = data.iloc[last_frame,1]\n    rate = (last_frame - first_frame)/(first_ttf - last_ttf)\n    print(\"Earthquake\", failure_index, \"frame range\", first_frame, last_frame, \"time range\", first_ttf, last_ttf, \"measured frame rate\", rate, \"Hz\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"56ae6c874a789e128d8d7315b93c72c614ca01f6"},"cell_type":"markdown","source":"I hope, that:\n* my research and produced outcome will be useful. \n* we'll get more insights from organizers shortly about the way training dataset was prepared and it's relation to test data."}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}