{"cells":[{"metadata":{"_uuid":"5c30ec08ee479e47fa34b4bfbe812f681757a742"},"cell_type":"markdown","source":"**The Basics**\n\nWe have a very large continuous csv of data.  There are 629 million rows of data.  Test data consists of snippets of 150,000 data points.  That's about 4194 full segments.  Based on the EDA  [here](http://https://www.kaggle.com/artgor/seismic-data-eda-and-baseline) it looks like there are 16 different failures in the data.  \n\n> The data within each test file is continuous, but the test files do not represent a continuous segment of the experiment; thus, the predictions cannot be assumed to follow the same regular pattern seen in the training file. \n\nThe test data, according to the description, is made up of random segments cut from a strip of data.  Basically, it's a warning that if you feed all the training data into the model straight, it may well pick up on some underlying pattern.\n\n"},{"metadata":{"_uuid":"fc01831d5afc31210672914822aad65c58ac86ea"},"cell_type":"markdown","source":"**First Questions**\n\nThe first question we need to answer is how to handle the occurance of a failure.  Looking at the EDA failures seem to reset the clock on time to failure (by letting stress out of the rock).  However they may also change to rate of time until the next failure.  If we were simply to iterate through the time series in batches of 150,000 approximately 0.4% of the data would have a fault occure during it?  Is this significant?  Probably not.\n\nThe next question, that would be helpful to answer, is what is a labratory earthquake?  Is this simulated data?  Earthquakes (very small) caused by fracking or some other form of explosion?\n\nIs all data in the test file from the same signal ID?  How many signal IDs are in each test file?"},{"metadata":{"_uuid":"4e6d4d53052cbbf2e696c6976652a6775c944e46"},"cell_type":"markdown","source":"**First Challenges**\nWe need to reduce the dimensionality of this data and create synthetic data poins.  \n\n1.  We need to be able to randomly pull a continous series of 150,000 data points (one entry) from the file.  Records the final time to failure datapoint.  Trash the rest of the time to failure data points, transpose the time rows to columns, and then append on the target (time to failure) datapoint.\n\n2.  This is a big data problem as the file to too large to load into memory all by itself.\n\n3.  We need to be able to reduce the dimensionality of the data.  While some forms of dimensionality reduction only require one row of data (for example computing its mean), others requires multiple rows.  This means we would need to to pull a statistically representative sample -- probably at least 500 of the new rows (1500 or everything would be better).  Create the model, export it and reimport it (so that we only need to crunch the numbers once).  Basically, create a PCA (or what have you) model once and resuse it every time a kernel is submitted.\n\n4.  The challenges with the above are big data (handling these very large rows), as well as how to export, save and import models.\n\n5.  Once we have a method for reducing the dimensionality of the data, we need to create a new data file with about 60,000 dimensionality reduced rows by randomly pulling batches of 150,000 data points and performing transforms on them.\n"},{"metadata":{"_uuid":"05e435dc6b6db4626ea1ae8cf7a685977ed83097"},"cell_type":"markdown","source":"**Modeling Challenges**\n\nOnce we have a new, dimensionality reduced, data file, we need to apply modeling.  The answer is almost, certainly, a DNN, though possibly in an ensemble with a statistical approach (SVM).  \n\nThis suggests that we might be looking at a two languages challenge.  Preprosessing data in R, then modeling it in Python/Keras."},{"metadata":{"_uuid":"e9eddf5a4ecb593f7ed5980a2cddcc2ff671b89c"},"cell_type":"markdown","source":"**Resources**\n\nThe following resources might be helpful.  \n\n[R TSrepr package](https://petolau.github.io/TSrepr-clustering-time-series-representations/)\n\n[Basic Info on Big Data with R](http://http://www.columbia.edu/~sjm2186/EPIC_R/EPIC_R_BigData.pdf)"},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","collapsed":true,"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":false},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}