{"cells":[{"metadata":{"_uuid":"ae74956485b850613b3fc2d0a4c0999e7ed1b19a"},"cell_type":"markdown","source":"# Introduction"},{"metadata":{"_uuid":"aed94d087b42350753f851f2fcfa86617c8ad96a"},"cell_type":"markdown","source":"I joined this competition only a **few days ago** and the first problem I faced has been the data loading. My kernels always crashed due to RAM limitations while loading \"parquet\" files. This notebook has a **simple aim**: summarising the competition data and **generate two sets** (train and test) **easily usable in other kernels**, mainly focused on modelling.\n\nThanks to [bluexleoxgreen](https://www.kaggle.com/bluexleoxgreen/simple-feature-lightgbm-baseline), especially for the idea of **splitting** the huge test set into subsets of **2K** columns.\n\nChanges from the previous version: I deleted the part where I collected samples to put them into the exported sets. The sampling rates where too low and a lot of information about peaks and errors were lost. Here, instead, I try to add the information about amplitudes and phases of the first harmonics, with a downsampling much more \"rich\" than the previous one.\n\nRunning times of the kernel: **15 minutes**.\n\nIn any case, I'd highly appreciate comments or suggestions to improve my Python or the efficiency of this notebook (better slicing, multithreading or whatever)."},{"metadata":{"_uuid":"c5b82a7cd3746c6e04bb9ead9f57ce79b357860c"},"cell_type":"markdown","source":"# Imports"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n% matplotlib inline","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ff84b9cfc47d635a269cf9473a8492241850dc68"},"cell_type":"code","source":"import warnings\nwarnings.filterwarnings('ignore')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"77a5d7c53ac56117456d7569004dbd252b0a16aa"},"cell_type":"code","source":"import gc\nimport pyarrow.parquet as pq","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4a7c236d6bdbd04530a0918c88f6e79dafee90ba"},"cell_type":"code","source":"import os\nprint(os.listdir(\"../input\"))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b56cf672c8010a4872f065e9d8ca9e0aa46489be"},"cell_type":"markdown","source":"# Data Loading"},{"metadata":{"trusted":true,"_uuid":"c43eaa24f5e08cb32a3aade1a873865055c0b77f"},"cell_type":"code","source":"metadata_train = pd.read_csv(\"../input/metadata_train.csv\")\nmetadata_train.info()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"38ac5c1552c3fb370f3eb5b7e0885af062ccac23"},"cell_type":"code","source":"metadata_train.head(12)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"5b44962a78f634a13a5d916c1e7feac66096e9c9"},"cell_type":"code","source":"metadata_test = pd.read_csv(\"../input/metadata_test.csv\")\nmetadata_test.info()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f7f635778f3e3e798d3e70d6beee03c945c408ac"},"cell_type":"code","source":"metadata_test.head(12)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"96e438a30274836d25785c66bc7483b0151c46ab"},"cell_type":"markdown","source":"The test set has about **3 times** the rows of the training set."},{"metadata":{"trusted":true,"_uuid":"9529b9777020c6ec5e80746579980e11fb275d26"},"cell_type":"code","source":"sample_submission = pd.read_csv(\"../input/sample_submission.csv\")\nsample_submission.info()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3beaff7c313fb1a3c54b635f9442ffdccd0716f4"},"cell_type":"code","source":"sample_submission.head(12)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"761dbd82a24e11dbeb69abd0a288f8685c1c36e8"},"cell_type":"markdown","source":"# Train Set Preparation (and a bit of EDA)"},{"metadata":{"_uuid":"683846390cf169f88d3455b637f061adb19dc35b"},"cell_type":"markdown","source":"Let's start loading the train.parquet file, which is actually a \"second level\" train set, containing information to put into the \"usual\" train set, metadata_train."},{"metadata":{"trusted":true,"_uuid":"45c5c93af021fcd78ab0e0cf911a4e2242cfed28"},"cell_type":"code","source":"%%time\ntrain = pq.read_pandas('../input/train.parquet').to_pandas()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"43b214eb23d28c530728b3b868363c5131f57229"},"cell_type":"code","source":"train.info()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"bc0341602a034850ffaed926984b9a95096d71da"},"cell_type":"code","source":"train.iloc[0:7,0:10]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b2951297487c0f32bc3cc1a4f88b8bf06e4d7d8d"},"cell_type":"markdown","source":"The columns are the FK (signal_id). The values are voltage levels. The rows are sampling points (one every 20ms). Let's look at one signal:"},{"metadata":{"trusted":true,"_uuid":"a21f2e11af46731e0554c2e7dd525a0f663f5c86"},"cell_type":"code","source":"x = train.index\ny = train.iloc[:,0]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ad405962ed07b1639fb72e57ebe0f11de55460cc"},"cell_type":"code","source":"fig = plt.figure(figsize=(12,4))\nax1 = fig.add_axes([0,0,1,1])\nax1.plot(x,y,color='lightblue')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b8ea555b15695bdfcc943a649963aac903ccf7c9"},"cell_type":"code","source":"fig = plt.figure(figsize=(12,4))\nax1 = fig.add_axes([0,0,1,1])\nax1.set_xlim([10,300])\nax1.set_ylim([10,30])\nax1.plot(x,y,marker='o', color='orange')\nax2 = fig.add_axes([0.7,0.7,0.2,0.2])\nax2.plot(x,y, color='lightblue')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b580be5bced13fb60cd586f560160c92f87966d9"},"cell_type":"markdown","source":"Now, let's put our attention on trios, starting from a \"good\" one."},{"metadata":{"trusted":true,"_uuid":"90074df9bcddca3c082eb9242245e1862bc47ba9"},"cell_type":"code","source":"x = train.index\ny0 = train.iloc[:,0]\ny1 = train.iloc[:,1]\ny2 = train.iloc[:,2]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d3d3332e65a899dd8cb76a5303c910b6cefd4c93"},"cell_type":"code","source":"fig = plt.figure(figsize=(12,4))\nax1 = fig.add_axes([0,0,1,1])\nax1.plot(x,y0,color='blue')\nax1.plot(x,y1,color='red')\nax1.plot(x,y2,color='green')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1b3079ca2cf642956a6143e475ce1bebfc888874"},"cell_type":"code","source":"np.mean(y0)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ed650b6276943ced4723aa79db0b3cc5533f42ef"},"cell_type":"code","source":"np.min(y0)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a9b36dc9116c92015f1293700fe32eea7ddb9f9c"},"cell_type":"code","source":"np.max(y0)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"65795c2ca6fb95aa7027594f9eaeaafec92ea08a"},"cell_type":"code","source":"np.std(y0)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a0d38a3f5ab9403d6026bc40405f71c0d9cdad92"},"cell_type":"markdown","source":"And now, here is one with target=1, i. e. a \"bad example\":"},{"metadata":{"trusted":true,"_uuid":"f6ffb840c380d7c6980dc1201aaaca05eca1404d"},"cell_type":"code","source":"y0 = train.iloc[:,3]\ny1 = train.iloc[:,4]\ny2 = train.iloc[:,5]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"7f5f1902097a582a06266f5833b8b07c80d5a5f9"},"cell_type":"code","source":"fig = plt.figure(figsize=(12,4))\nax1 = fig.add_axes([0,0,1,1])\nax1.plot(x,y0,color='blue')\nax1.plot(x,y1,color='red')\nax1.plot(x,y2,color='green')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"12172ad7a6146127f482d18f3510d43a022b5070"},"cell_type":"markdown","source":"I'd say that a failure is linked to peaks in two (or more phases), but let's go on. "},{"metadata":{"trusted":true,"_uuid":"4157d3dd149a263067f9aa15792914ed018dc4e4"},"cell_type":"code","source":"np.mean(y0)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b3791d822beaa23089277ea4694980d07a4990cf"},"cell_type":"code","source":"np.min(y0)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9e8169fbe7fe79f305af981220115cec7f3925ec"},"cell_type":"code","source":"np.max(y0)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"cc51fa9fe38ffdf3933f2ec60c872f65d2ca52a6"},"cell_type":"code","source":"np.std(y0)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8ca9b67d3880322b196a1bffbbb0b2e1317506b6"},"cell_type":"markdown","source":"At this point we want to add some common indexes such as mean, std etc ... "},{"metadata":{"trusted":true,"_uuid":"1f139829016a2d80a01cf702e5fba29a1fe4efe5"},"cell_type":"code","source":"row_nr = train.shape[0]\nrow_nr","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"72bb172c6ceb9e92597e21a0e069339ae04fdad2"},"cell_type":"code","source":"index_group_size=100","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"31c9dfa411f9d4513b7f315271b2ffc5fb8da873"},"cell_type":"code","source":"time_sample_idx=np.arange(0,row_nr,index_group_size)\ntime_sample_idx[0:10]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"df25d3bd40b45b10bb071b05c1534ad03301516d"},"cell_type":"code","source":"train_down=train.iloc[time_sample_idx,:]\ntrain_down.iloc[:,0:10].head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"101a3cfdd786c758dae4728db9d207fd9ca4cdbf"},"cell_type":"code","source":"import numpy.fft as ft","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9a8c4717df7346503544d1ff2c78db7f3b808081"},"cell_type":"code","source":"def Amplitude(z):\n    return np.abs(z)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f81cf5d6d39542fb2dc1853c2319b243e4a1b9d0"},"cell_type":"code","source":"def Phase(z):\n    return (np.arctan2(z.imag,z.real))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"818a988df271948aa257223c72d3aae57597be8c"},"cell_type":"code","source":"df_harm=pd.DataFrame()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"500471f0c6c62fa1a100c15512479981bd1bcb89"},"cell_type":"code","source":"def find_dfa(df_source, df_dest,num_harm,base_col):\n    # init\n    dfa=df_dest.iloc[:,0:base_col]\n    num_ap_cols = int(num_harm/2)\n    for j in range(0,num_ap_cols) :\n        dfa['Amp'+str(j)] = 0\n        dfa['Pha'+str(j)] = 0\n    dfa['ErrFun'] = 0\n    dfa['ErrGen'] = 0\n    # calc\n    for i in range(0,len(df_source.columns)):\n        dfa.loc[i]=0\n        s=df_source.iloc[:,base_col+i]\n        SF=ft.rfft(s)\n        SF_Fundam=np.zeros(SF.size, dtype=np.complex_)\n        SF_Filtered=np.zeros(SF.size, dtype=np.complex_)\n        SF_Fundam[0:2]=SF[0:2]\n        SF_Filtered[0:num_harm]=SF[0:num_harm]\n        s_fun_rec=ft.irfft(SF_Fundam)\n        s_gen_rec=ft.irfft(SF_Filtered)\n        for j in range(0,num_ap_cols):\n            dfa.iloc[i,base_col+2*j] = Amplitude(SF_Filtered[j])\n            dfa.iloc[i,base_col+2*j+1] = Phase(SF_Filtered[j])\n        dfa.iloc[i,base_col+2*num_ap_cols] = np.sqrt(np.mean((s-s_fun_rec)**2))\n        dfa.iloc[i,base_col+2*num_ap_cols+1] = np.sqrt(np.mean((s-s_gen_rec)**2))\n    return dfa","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e6e929a9cfff675f6a86c8dea0b81b19fc263b26"},"cell_type":"code","source":"%%time\ntrain_max = train.apply(np.max)\ntrain_min = train.apply(np.min)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"22d4092ecf63ad8d3b101ac425376f6ef85f6b34"},"cell_type":"code","source":"%%time\ntrain_mean = train_down.apply(np.mean)\ntrain_std = train_down.apply(np.std)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a07105c10aae06d316101184c0e0d166cee655a8"},"cell_type":"code","source":"%%time\ndf_harm=pd.DataFrame()\nnum_harm=10\ndf_harm=find_dfa(train_down,df_harm,num_harm,0)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4ea5dd32e30e0a1cb0c3b3821641f1aed4a58e6c"},"cell_type":"code","source":"df_harm.iloc[:,0:10].head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d6d82a159c7d795ac483e546e62e0d03aed91d1b"},"cell_type":"code","source":"metadata_train['mean']=train_mean.values\nmetadata_train['max']=train_max.values\nmetadata_train['min']=train_min.values\nmetadata_train['std']=train_std.values\nfor j in range(0,int(num_harm/2)) :\n    metadata_train['Amp'+str(j)] = df_harm['Amp'+str(j)]\n    metadata_train['Pha'+str(j)] = df_harm['Pha'+str(j)]\nmetadata_train['ErrFun'] = df_harm['ErrFun']\nmetadata_train['ErrGen'] = df_harm['ErrGen']","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1e044e7a089795b179afbc3f75f5064390bd0265"},"cell_type":"code","source":"metadata_train.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1446003fa6f33698a0e949761acdc0870e8ca652"},"cell_type":"code","source":"df_train = metadata_train\ndf_train.to_csv('df_train.csv', index=False)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"07e3f024417ba9432f7ea1c5b8be758d7688b9f2"},"cell_type":"markdown","source":"# Test Set Preparation"},{"metadata":{"_uuid":"c0a1772451a221ab112a35603ced7e8f1bbbd7b7"},"cell_type":"markdown","source":"It's impossibile to load the test set as a whole, let's define a chunk size:"},{"metadata":{"trusted":true,"_uuid":"168b8c6aa5b4fbfa9ab2e7003cf1767d2468f1be"},"cell_type":"code","source":"col_group_size=2000","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"05f97712abed8ad787704fd8c333a7093223a5b4"},"cell_type":"markdown","source":"In addition, let's free all the avalable resources:"},{"metadata":{"trusted":true,"_uuid":"e3791d9d933ee4a5fa0bf08b511fe55870c41425"},"cell_type":"code","source":"gc.collect()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"89b9ba700c035ffd8dffe3524943f87ef67754b0"},"cell_type":"markdown","source":"Now let's restart from the metadata_test:"},{"metadata":{"trusted":true,"_uuid":"aba322bca11175b5676264cab878cd6faf2662cc"},"cell_type":"code","source":"metadata_test = pd.read_csv(\"../input/metadata_test.csv\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"8fb5a5f08633d55045c2987ee0d09a3f5b8b2fd1"},"cell_type":"code","source":"metadata_test['target']=-1\nmetadata_test['mean']=0\nmetadata_test['max']=0\nmetadata_test['min']=0\nmetadata_test['std']=0\nfor j in range(0,int(num_harm/2)) :\n    metadata_test['Amp'+str(j)] = 0\n    metadata_test['Pha'+str(j)] = 0\nmetadata_test['ErrFun'] = 0\nmetadata_test['ErrGen'] = 0","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"84faa391fa3694b92a16004d83e33e7ad611131e"},"cell_type":"code","source":"metadata_test.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c73b2f79800a784d85f9938f8b12fc82ef90de2d"},"cell_type":"code","source":"metadata_test.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7ffce4e7f8deb10714fa0452e5f303b067001915"},"cell_type":"markdown","source":"Here is the function definition:"},{"metadata":{"trusted":true,"_uuid":"f8610e40d037b4ea187fb27af34eb89dff012087"},"cell_type":"code","source":"def add_info_test(metadata_df,time_sample_idx_1,col_group_size):\n    col_id_start_0=np.min(metadata_test['signal_id'])\n    col_id_start=col_id_start_0\n    col_id_last=np.max(metadata_test['signal_id'])+1\n    n_groups = int(np.round((col_id_last-col_id_start)/col_group_size))\n    print('Steps = {}'.format(n_groups))\n    for i in range(0,n_groups):\n        col_id_stop = np.minimum(col_id_start+col_group_size,col_id_last)\n        col_numbers = np.arange(col_id_start,col_id_stop)\n        print('Step {s} - cols = [{a},{b})'.format(s=i,a=col_id_start,b=col_id_stop))\n        print('   Adding Stats...',end=\"\")\n        col_names = [str(col_numbers[j]) for j in range(0,len(col_numbers))]\n        test_i = pq.read_pandas('../input/test.parquet',columns=col_names).to_pandas()\n        test_i_d1=test_i.iloc[time_sample_idx_1,:]\n        test_mean_i = test_i_d1.apply(np.mean)\n        test_max_i  = test_i.apply(np.max)\n        test_min_i  = test_i.apply(np.min)\n        test_std_i  = test_i_d1.apply(np.std)\n        r_start = col_id_start - col_id_start_0\n        r_stop = r_start + (col_id_stop-col_id_start)\n        metadata_df.iloc[r_start:r_stop,4] = test_mean_i[0:col_id_stop-col_id_start].values\n        metadata_df.iloc[r_start:r_stop,5] = test_max_i[0:col_id_stop-col_id_start].values\n        metadata_df.iloc[r_start:r_stop,6] = test_min_i[0:col_id_stop-col_id_start].values\n        metadata_df.iloc[r_start:r_stop,7] = test_std_i[0:col_id_stop-col_id_start].values\n        print('   Adding FFT...')\n        df_harm=pd.DataFrame()\n        df_harm=find_dfa(test_i_d1,df_harm,10,0)\n        num_ap_cols = int(num_harm/2)\n        fft_base_col=8\n        for j in range(0, num_ap_cols) :\n            metadata_df.iloc[r_start:r_stop,fft_base_col+2*j] = df_harm.iloc[r_start:r_stop,2*j]\n            metadata_df.iloc[r_start:r_stop,fft_base_col+2*j+1] = df_harm.iloc[r_start:r_stop,2*j+1]\n        metadata_df.iloc[r_start:r_stop,fft_base_col+num_harm] = df_harm.iloc[r_start:r_stop, num_harm]\n        metadata_df.iloc[r_start:r_stop,fft_base_col+num_harm+1] = df_harm.iloc[r_start:r_stop, num_harm+1]\n        col_id_start=col_id_stop\n    return (metadata_df)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5dbcaeaac51b1f3c6f6a5a4b18ee8dd7bd386ad2"},"cell_type":"markdown","source":"And this is its call:"},{"metadata":{"trusted":true,"_uuid":"0a7550debfbfde5740cf4f6e791084b2391293ab"},"cell_type":"code","source":"%%time\nmetadata_test1=add_info_test(metadata_test,time_sample_idx,col_group_size)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9aefa00cf5f002caefe1be2626833e23379b3043"},"cell_type":"code","source":"metadata_test1.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ecf4b232d6a83b816001781178b1b9af77e7a431"},"cell_type":"code","source":"metadata_test1.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"44bd6cb1b4bba4d833c72ffc888fa933a2da3cb5"},"cell_type":"code","source":"metadata_test1.iloc[0:12,:]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6ef5514c4da7eea921164f69a7a29f942e3046fc"},"cell_type":"markdown","source":"So we have, finally, our exportable test set:"},{"metadata":{"trusted":true,"_uuid":"f5ed727bb936f9f6a65847ee06cb5162878eae15"},"cell_type":"code","source":"metadata_test1.to_csv('df_test.csv', index=False)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"805100ad22081b51b44e42c9f8b3004f6e4b355f"},"cell_type":"markdown","source":"Now it's time to have fun with a new notebook based on these data and focused on modelling. \n\nLike said Pink some years ago, \"[Let's get the party started!](https://www.youtube.com/watch?v=QRINgISPUWQ)\""}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}