{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load in \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the \"../input/\" directory.\n# For example, running this (by clicking run or pressing Shift+Enter) will list the files in the input directory\n\nimport os\nprint(os.listdir(\"../input\"))\n\n# Any results you write to the current directory are saved as output.","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"be3ec9281dd66c8972c0bff00c005af1befa9534"},"cell_type":"markdown","source":"The goal of this kernel is to **establish a baseline **and **explain everything step by step** so anyone can get started. Code to read in and process the data is based on this kernel: <a href='https://www.kaggle.com/inversion/basic-feature-benchmark'>Basic Feature Benchmark.</a>\n\nAny questions, comments, or suggestions are always welcome."},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"#Read in the training data\ntrain_df = pd.read_csv('../input/train.csv',\n                    dtype={'acoustic_data': np.int16,\n                           'time_to_failure': np.float64}) ","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f02d74f43c37f0512b454b6bdb05e4eeed032fec"},"cell_type":"markdown","source":"By specifying the datatypes, pandas can read in the csv faster and use less memory. Without specifying the datatypes the kernel can crash because the file is so large"},{"metadata":{"trusted":true,"_uuid":"c7a1369930a9d4dd57ee78391d30468db085c8f2"},"cell_type":"code","source":"#Look at the data and realize its not like a normal dataset\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3305a4768e54f284388c7499068b3b5694360667"},"cell_type":"code","source":"train_df.shape","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"de7a99f76dc42334bc6af56c8c0c32d64872b126"},"cell_type":"markdown","source":"Instead of each row being one sample, we have one long continous set of data. **We will split this one long sample into many different samples.**\n\nThe first column*** 'acoustic data'***  is our only feature. We use that feature to predict ***'time_to_failure'*.**\n\n**Each new sample will be created from 150,000 rows** of only the *'acoustic_data'* column. The target variable is the* 'time_to_failure' *at the last row of the sample. (We use 150,000 rows because thats how long the test samples are)"},{"metadata":{"trusted":true,"_uuid":"9a5e023664563ea393fca2d7e10d873c988d2f25"},"cell_type":"code","source":"#Define how long we want each sample to be\nsample_length = 150000\n\n#Divide length of our dataframe by how long we want each sample\nnum_samples = int((len(train_df) / sample_length)) ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"333c92b6fad245d4c4d34a5d4f4cb7b0b92d0d2b"},"cell_type":"code","source":"#This is how many samples we will create\nnum_samples","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"64ebf37e7c140eca81b874e30e03ea7ad9886326"},"cell_type":"code","source":"#This is a list of features we will create\ncols = ['mean','median','std','max',\n        'min','var','ptp','10p',\n        '25p','50p','75p','90p']","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d15c6cb60f180763f0485f5b9ab39fabf25f7f40"},"cell_type":"markdown","source":"Our goal is to condense each sample into one row . **We create these features to extract information out of the 150,000 rows that make up each sample. After we extract those features, we can then represent each sample as one row in a dataframe.**"},{"metadata":{"trusted":true,"_uuid":"f0ae2e7d4f391a5ce4d607d67b82fab50433aafc"},"cell_type":"code","source":"#This creates an empty dataframe for now\n#Later we will fill it with values\nX_train = pd.DataFrame(index=range(num_samples), #The index will be each of our new samples\n                       dtype=np.float64, #Assign a datatype\n                       columns=cols) #The columns will be the features we listed above","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"62749f7cdb225346b48bf9dbcb87dfd3b24eca12"},"cell_type":"code","source":"#This creates a dataframe for our target variable 'time_to_failure'\ny_train = pd.DataFrame(index=range(num_samples),\n                       dtype=np.float64, \n                       columns=['time_to_failure']) #Our target variable","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"cd1fe66bf1182810aa9a7b247e366de6f27aa641"},"cell_type":"code","source":"#Now we create the samples\nfor i in range(num_samples):\n    \n    #i*sample_length = the starting index (from train_df) of the sample we create\n    #i*sample_length + sample_length = the ending index (from train_df)\n    sample = train_df.iloc[i*sample_length:i*sample_length+sample_length]\n    \n    #Converts to numpy array\n    x = sample['acoustic_data'].values\n    \n    #Grabs the final 'time_to_failure' value\n    y = sample['time_to_failure'].values[-1]\n    y_train.loc[i, 'time_to_failure'] = y\n    \n   #For every 150,000 rows, we make these calculations\n    X_train.loc[i, 'mean'] = np.mean(x)\n    X_train.loc[i, 'median'] = np.median(x)\n    X_train.loc[i, 'std'] = np.std(x)\n    X_train.loc[i, 'max'] = np.max(x)\n    X_train.loc[i, 'min'] = np.min(x)\n    X_train.loc[i, 'var'] = np.var(x)\n    X_train.loc[i, 'ptp'] = np.ptp(x) #Peak-to-peak is like range\n    X_train.loc[i, '10p'] = np.percentile(x,q=10) \n    X_train.loc[i, '25p'] = np.percentile(x,q=25) #We can also grab percentiles\n    X_train.loc[i, '50p'] = np.percentile(x,q=50)\n    X_train.loc[i, '75p'] = np.percentile(x,q=75)\n    X_train.loc[i, '90p'] = np.percentile(x,q=90)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9145d091948d91c2ae35ab75aa3952384fb46cd8"},"cell_type":"code","source":"#Creates a simple train, test split\nfrom sklearn.model_selection import train_test_split\n\nX_train, X_val, y_train, y_val = train_test_split(X_train, y_train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"edcb5dca68c246ce3de0e2cfecaa705430b66ef7","_kg_hide-output":true},"cell_type":"code","source":"#Fit a random forest\nfrom sklearn.ensemble import RandomForestRegressor\n\n#This creates the Randomforest with the given parameters\nrf = RandomForestRegressor(n_estimators=100, #100 trees (Default of 10 is too small)\n                          max_features=0.5, #Max number of features each tree can use \n                          min_samples_leaf=30, #Min amount of samples in each leaf\n                          random_state=11)\n\n#This trains the random forest on our training data\nrf.fit(X_train,y_train)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2549173c76d7d9922c7e4a01f7ea48a62d98998f"},"cell_type":"markdown","source":"**n_estimators**: the number of decision trees in our random forest. \n\n**max_features**: The max amount of features each decision tree can use. This can be an integer (max_features=10 means use 10 features) or a ratio (0.5 means each tree is fit using half of the original features. Features are selected at random) \n\n**min_samples_leaf**: Minimum number of samples in each leaf of the decision trees"},{"metadata":{"trusted":true,"_uuid":"e480ca0f9d67bc36917981ad7a2839b57e3a0143"},"cell_type":"code","source":"#Score the model\nfrom sklearn.metrics import mean_absolute_error\n\nmean_absolute_error(y_val, rf.predict(X_val))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f206e6558ce40669561f6ab5cc08fa56baaf2eaf"},"cell_type":"markdown","source":"Now that you have a baseline, here are some next steps to explore:\n\n* Try to think of any other useful features you can extract for each sample. Here are more summary statics to try from <a href='https://docs.scipy.org/doc/scipy/reference/stats.html#summary-statistics'>SciPy</a>. SciPy also has different signal processing functions to check out.\n\n* Calculate the min, max, mean, etc over smaller chunks of the 150,000 rows per sample and add those as more features.\n\n* Create more samples. Instead of our training data being generated from seperate sections of 150,000 rows with no overlap, you can adjust the indices so that the samples overlap. If you do this be very careful to make sure no samples in your validation set overlap with your training data.\n\n* Try a LightGBM, XGBoost, or NN model.\n\n* Plot some samples to explore the data\n\n* Plot feature importances to see what your model finds important."},{"metadata":{"_uuid":"8326f08925beb734baffdd15be111714e2ada286"},"cell_type":"markdown","source":"How to submit predictions from the kernel:"},{"metadata":{"trusted":true,"_uuid":"0df753878db33295d6f6ec13671a801e17222004"},"cell_type":"code","source":"#Read in the sample submission. We can use that as a dataframe to grab the segment ids\nsubmission = pd.read_csv('../input/sample_submission.csv',\n                         index_col = 'seg_id')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"2d0dffcf95b34df081ba39256e62a90cc4142bea"},"cell_type":"code","source":"submission.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"bacd432df617252635bea88c0fae58d9e635cdcb"},"cell_type":"code","source":"#Creates a test dataframe\nX_test = pd.DataFrame(columns=X_train.columns, #Use the same columns as our X_train\n                      dtype=np.float64,\n                      index=submission.index) #Use the index ('seg_id') from the sample submission","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ce16b7c675c62446e355038ab5a418eae80cfd15"},"cell_type":"markdown","source":"In the test folder each segment is its own csv file. To eactract the features we loop through each segment id one at a time."},{"metadata":{"trusted":true,"_uuid":"d9d5eeec8bb9db68a542961dbed437713b0433b5"},"cell_type":"code","source":"for i in X_test.index:\n    \n    #Read in that segments csv file\n    #By putting f before the string we can put any values between {} and it will be treated as a string\n    seg = pd.read_csv(f'../input/test/{i}.csv') \n                                            \n    #Grab the acoustic_data values\n    x = seg['acoustic_data'].values\n\n    #These are the same features we calcuted on the training data\n    X_test.loc[i, 'mean'] = np.mean(x)\n    X_test.loc[i, 'median'] = np.median(x)\n    X_test.loc[i, 'std'] = np.std(x)\n    X_test.loc[i, 'max'] = np.max(x)\n    X_test.loc[i, 'min'] = np.min(x)\n    X_test.loc[i, 'var'] = np.var(x)\n    X_test.loc[i, 'ptp'] = np.ptp(x)\n    X_test.loc[i, '10p'] = np.percentile(x,q=10) \n    X_test.loc[i, '25p'] = np.percentile(x,q=25)\n    X_test.loc[i, '50p'] = np.percentile(x,q=50)\n    X_test.loc[i, '75p'] = np.percentile(x,q=75)\n    X_test.loc[i, '90p'] = np.percentile(x,q=90)\n    ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3807f8c7a6d326fa8a553decf078ad7bfd8fd35e"},"cell_type":"code","source":"#Predict on the test data\ntest_predictions = rf.predict(X_test)\n\n#Assign the target column in our submission to be our predictions\nsubmission['time_to_failure'] = test_predictions","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"aaf30b81f5542ac1cbc78f3b187b9b805c0a160f"},"cell_type":"code","source":"#Output the predictions to a csv file\nsubmission.to_csv('submission.csv')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4b0a1e5c47009e497ed179c317267086a5753699"},"cell_type":"markdown","source":"Now click Commit. After it runs, click Open Version. Scroll down to output and click submit to competition."}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}