{"cells":[{"metadata":{"_uuid":"59e7874ebb21d890b70e75c7b7fc40419d15d3d9"},"cell_type":"markdown","source":"This kernel creates a vector row_breaks, which marks every point in the dataset where TTF takes a jump in measurement due to:\n\n1 - in between bins:\nBecause the measurement tool has an artifact, after every bin TTF is missing some microseconds,\nwe can use this to our advantage of finding out where exactly the bins end. It seems the bins are not the same size as the 150.000 sized test-bins, \nwhich might have an impact on model accuracy. The bins rather seem to be 151550+- in size.\nSo I have made this kernel to exactly point out where the bins start / end.\n\n2 - a TTF reset after a quake happened:\nAlso, there are 16 quakes in this dataset. Which also means that TTF will be reset 16 times,\nas the data is contiuous, these resets take place inside a bin rather than at start or end.\nSo it's in our interest to also point out where TTF resets, and add this to our row_breaks vector."},{"metadata":{"_uuid":"520bb91fb44182bc50c21711700387df5f2875d4","_execution_state":"idle","trusted":true,"scrolled":true},"cell_type":"code","source":"#----------------- Read in Data -------------------#\n\nlibrary(data.table) #for fread\ntrain <- fread('../input/train.csv')\n\n#----------------- END Read in Data -------------------#","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3527d6a92564edb7914c1906bc75bea7e9b5facf"},"cell_type":"markdown","source":"We will loop up to 4200 times, that is a little over our estimated number of bins.\nFor every bin we will do a smaller loop which checks 200 datapoints (from 151400 to 151600 rows) - to check where the jump of TTF is."},{"metadata":{"trusted":true,"_uuid":"cd89917e63727bdc637f9c9fe9884f7990012320"},"cell_type":"code","source":"#----------------- Create breaks%1 -------------------#\n\nrow_breaks <- vector(\"list\") ; row_breaks <- append(row_breaks,1)\n\noffset <- 151400\ncurrentrow <- 151400 # to find the first end-of-bin, but will be increased with offset every run\n\nfor(o in c(1:4200)){\n    # print(round(o/42,2)) # prints in %-points how far we are in the loop\n    for (i in c(1:200)){\n    currentrow <- currentrow+1\n    previousrow <- currentrow-1\n    \n    if(train[previousrow,\"time_to_failure\"] - train[currentrow,\"time_to_failure\"]>0.0001)\n    {\n      row_breaks <<- append(row_breaks, currentrow)\n      break\n    }\n    \n    #statement to stop if we are at the end of the data:\n    if(is.na(train$acoustic_data[currentrow])==TRUE){\n      break\n    }\n  }\n  currentrow <- currentrow + offset\n  if(is.na(train$acoustic_data[currentrow])==TRUE){\n    break\n  }\n}\n\n#----------------- END breaks%1 -------------------#\n\n\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b52ecf20fdd8a249d8dec404373601f61c29e744"},"cell_type":"markdown","source":"Now we are going to check where TTF comepletely resets after a quake.\nThis might be important, as after every quake you will have one bin which has really lowsttf (near 0) + at the same time really high ttf's.\nThis in turn might make our model training somewhat more cloudy, since this only affects 16 bins - we prefer to know where these breaks are,\nso that we can later remove these bins.\n\nSince this took about +-40 minutes to run on a beefy CPU,\nI've put the results of the code below."},{"metadata":{"trusted":true,"_uuid":"4a1fb52693f71f3ef3231ad8475bba93139f482d"},"cell_type":"code","source":"#----------------- Create Breaks%2 -------------------#\n\n# for (i in c(2:nrow(train))){\n#   currentrow <- i\n#   previousrow <- currentrow-1\n#   if((train$time_to_failure[currentrow]-train$time_to_failure[previousrow]) > 1){\n#     print(currentrow)\n#     print(train[currentrow])\n#   }\n# }\n\n# Results:\nttf_resets <- c(5656575,50085879,104677357,138772454,187641821,\n                  218652631,245829586,307838918,338276288,375377849,\n                  419368881,461811624,495800226,528777116,585568145,\n                  621985674)\n\n#----------------- END Create Breaks%2 -------------------#","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"66e12610484ab24aec5238b302de1639bd970a3c"},"cell_type":"markdown","source":"Now we need to append the 4152 bin breaks and the 16 ttf-reset breaks,\nand order them for ease-of use later:"},{"metadata":{"trusted":true,"_uuid":"fb91d647554fbfd747022c6a5c7a45ad258cc50d"},"cell_type":"code","source":"#----------------- Append Breaks -------------------#\n\nrow_breaks <- append(row_breaks,ttf_resets)\n\n# order row_breaks:\nrow_breaks <- sort(unlist(row_breaks), decreasing = FALSE)\n\n#----------------- END Append Breaks -------------------#","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"00567664fa64e209d92c1d814dff8c5e070a1c4e"},"cell_type":"markdown","source":"Next step would be creating the vars that you want, loop them over the different bins by using row_breaks\nand after you can delete the bins you don't like (for example the 32/partial sized bins)."}],"metadata":{"kernelspec":{"display_name":"R","language":"R","name":"ir"},"language_info":{"mimetype":"text/x-r-source","name":"R","pygments_lexer":"r","version":"3.4.2","file_extension":".r","codemirror_mode":"r"}},"nbformat":4,"nbformat_minor":1}