{"cells":[{"metadata":{"_uuid":"6df5acabcd9e823cf308353f2746513ea9230842"},"cell_type":"markdown","source":"# Preprocessing with Multiprocessing\n\nThis kernel uses the Python multiprocessing module to make use of all 4 cores you get in a Kaggle CPU kernel. The input data is split into four chunks, and each is processed using the given feature extraction function in parallel. The results are stored on disk in both scaled and non-scaled format. Output compression is used the avoid kernel output size limits.\n\nKaggle CPU kernels have 4 CPU cores, while GPU kernels have only 2 CPU cores. Running pre-processing in a separate kernel like this helps use both kernel types more optimally. CPU kernel with more cores to pre-process, GPU kernel to build and experiment with models using the CPU kernel output as data source.\n\nThe features in this version base on the ones in https://www.kaggle.com/braquino/5-fold-lstm-attention-fully-commented-0-694, and a few lag/diff ones I played with. Should be simple to tune for any other features/processing.\n\nAll the features are created in function \"summarize_df_np\". Change that to produce different features. \n\nThis kernel produces a set of output files as follows:\n\n- my_train.csv.gz: The raw data as processed (features, buckets, whatever you call them). Each signal separately in 160 rows per signal. \"raw\" as in not scaled.\n- my_train_scaled.csv.gz: The same data as my_train.csv.gz, scaled using min-max-scaler with -1 to 1 scale.\n- my_train_combined_scaled.csv.gz: The scaled data but all 3 signals per measurement on a single row.\n- my_test.csv.gz: Same as above but for test data.\n- my_test_scaled.csv.gz: Same as above but for test data.\n- my_test_combined_scaled.csv.gz: Same as above but for test data.\n"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import pandas as pd\nimport pyarrow.parquet as pq\nimport os\nimport numpy as np\nimport matplotlib.pyplot as plt","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"train_meta = pd.read_csv(\"../input/metadata_train.csv\")\n#train_meta.head(6)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"edb4e7d27240aa2e42d5d7222709bddea163a110"},"cell_type":"code","source":"test_meta = pd.read_csv(\"../input/metadata_test.csv\")\n#test_meta.head(6)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2c2ea5fb478f7bedd6b8a56bdf93108750f1d8bc"},"cell_type":"markdown","source":"## Feature Generation"},{"metadata":{"_uuid":"923f266a6fa5b79e879cef182c54c4333883c8a0"},"cell_type":"markdown","source":"The function to process all the chunks and generate features in the parallel running processes:"},{"metadata":{"trusted":true,"_uuid":"ff1200ea3537de65c81a826c3a60ab360b21d6e9"},"cell_type":"code","source":"#I use bkt as short for bucket, rather than bin for bin since I tend to read bin as binary\nbkt_count = 160\ndata_size = 800000\nbkt_size = int(data_size/bkt_count)\n\ndef summarize_df_np(meta_df, data_type, p_id):\n    count = 0\n    measure_rows = []\n\n    for measurement_id in meta_df[\"id_measurement\"].unique():\n        count += 1\n        idx1 = measurement_id * 3\n        input_col_names = [str(idx1), str(idx1+1), str(idx1+2)]\n        df_sig = pq.read_pandas('../input/'+data_type+'.parquet', columns=input_col_names).to_pandas()\n        df_sig = df_sig.clip(upper=127, lower=-127)\n\n        df_diff = pd.DataFrame()\n        for col in input_col_names:\n            df_diff[col] = df_sig[col].diff().abs()\n        \n        data_measure = df_sig.values\n        data_diffs = df_diff.values\n        sig_rows = []\n        sig_ts_rows = []\n        for sig in range(0, 3):\n            #take the data for each 3 signals in a measure separately\n            data_sig = data_measure[:,sig]\n            data_diff = data_diffs[:,sig]\n            bkt_rows = []\n            diff_avg = np.nanmean(data_diff)\n            for i in range(0, data_size, bkt_size):\n                # cut data to bkt_size (bucket size)\n                bkt_data_raw = data_sig[i:i + bkt_size]\n                bkt_avg_raw = bkt_data_raw.mean() #1\n                bkt_sum_raw = bkt_data_raw.sum() #1\n                bkt_std_raw = bkt_data_raw.std() #1\n                bkt_std_top = bkt_avg_raw + bkt_std_raw #1\n                bkt_std_bot = bkt_avg_raw - bkt_std_raw #1\n\n                bkt_percentiles = np.percentile(bkt_data_raw, [0, 1, 25, 50, 75, 99, 100]) #7\n                bkt_range = bkt_percentiles[-1] - bkt_percentiles[0] #1\n                bkt_rel_perc = bkt_percentiles - bkt_avg_raw #7\n\n                bkt_data_diff = data_diff[i:i + bkt_size]\n                bkt_avg_diff = np.nanmean(bkt_data_diff) #1\n                bkt_sum_diff = np.nansum(bkt_data_diff) #1\n                bkt_std_diff = np.nanstd(bkt_data_diff) #1\n                bkt_min_diff = np.nanmin(bkt_data_diff) #1\n                bkt_max_diff = np.nanmax(bkt_data_diff) #1\n\n                raw_features = np.asarray([bkt_avg_raw, bkt_std_raw, bkt_std_top, bkt_std_bot, bkt_range])\n                diff_features = np.asarray([bkt_avg_diff, bkt_std_diff, bkt_sum_diff])\n                bkt_row = np.concatenate([raw_features, diff_features, bkt_percentiles, bkt_rel_perc])\n                bkt_rows.append(bkt_row)\n            sig_rows.extend(bkt_rows)\n        measure_rows.extend(sig_rows)\n    df_sum = pd.DataFrame(measure_rows)\n    #df_sum = df_sum.astype(\"float32\")\n    return df_sum\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c44f90078690a7b2b890c28aa98551f2756c9577"},"cell_type":"markdown","source":"Function process_subtrain() is passed to the Python multiprocessing for all four cores/chunks. \n\nIt calls the feature processing function for the training data:"},{"metadata":{"trusted":true,"_uuid":"9634da7f8a175ce994422ab020034495d911e923"},"cell_type":"code","source":"def process_subtrain(arg_tuple):\n    meta, idx = arg_tuple\n    df_sum = summarize_df_np(meta, \"train\", idx)\n    return idx, df_sum","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"237a4e2528f52a0325b4dbe092cc00aacda80880"},"cell_type":"markdown","source":"The scaler to produce the scaled version on of the data once all four parallel chunks have finished processing:"},{"metadata":{"trusted":true,"_uuid":"5eb171550290426b46d947d0b95dc0aa37f42e59"},"cell_type":"code","source":"from sklearn.preprocessing import MinMaxScaler\n\nminmax = MinMaxScaler(feature_range=(-1,1))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2f4aa2a4de7ca6191fc7ff8405fdc5c41d228623"},"cell_type":"markdown","source":"## Multiprocessing"},{"metadata":{"_uuid":"e49916943038edd53f6479ab2c2f23c659a2e41f"},"cell_type":"markdown","source":"Function to create the chunks sizes/indices to split the data into chunks. Used for both train and test data:"},{"metadata":{"trusted":true,"_uuid":"9f1aa737d1d3b6c89545d720a7728758184d663f"},"cell_type":"code","source":"def create_chunk_indices(meta_df, chunk_idx, chunk_size):\n    start_idx = chunk_idx * chunk_size\n    end_idx = start_idx + chunk_size\n    meta_chunk = meta_df[start_idx:end_idx]\n    print(\"start/end \"+str(chunk_idx+1)+\":\" + str(start_idx) + \",\" + str(end_idx))\n    print(len(meta_chunk))\n    #chunk_idx in return value is used to sort the processed chunks back into original order,\n    return (meta_chunk, chunk_idx)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"83f3f7ba710ec9bfc31c914283da2d4798f5d5ca"},"cell_type":"markdown","source":"## Training dataset processing"},{"metadata":{"_uuid":"6f3b65d1ddcc7ac141ab9bb4442efb3b56133888"},"cell_type":"markdown","source":"Actual code to call multiprocessing for the training data:"},{"metadata":{"trusted":true,"_uuid":"e36caaae2709e45ef7fb50e16576fc7d21dbab9e"},"cell_type":"code","source":"from multiprocessing import Pool\n\nnum_cores = 4\n\ndef process_train():\n    #splitting here by measurement id's to get all signals for a measurement into single chunk\n    measurement_ids = train_meta[\"id_measurement\"].unique()\n    df_split = np.array_split(measurement_ids, num_cores)\n    chunk_size = len(df_split[0]) * 3\n    \n    chunk1 = create_chunk_indices(train_meta, 0, chunk_size)\n    chunk2 = create_chunk_indices(train_meta, 1, chunk_size)\n    chunk3 = create_chunk_indices(train_meta, 2, chunk_size)\n    chunk4 = create_chunk_indices(train_meta, 3, chunk_size)\n\n    #list of items for multiprocessing, 4 since using 4 cores\n    all_chunks = [chunk1, chunk2, chunk3, chunk4]\n    \n    pool = Pool(num_cores)\n    #this starts the (four) parallel processes and collects their results\n    #-> process_subtrain() is called concurrently with each item in all_chunks \n    result = pool.map(process_subtrain, all_chunks)\n    #parallel processing can be non-deterministic in timing, so here I sort results by their chunk id\n    #to maintain results in same order as in original files (to match metadata from other file)\n    print(\"sorting\")\n    result = sorted(result, key=lambda tup: tup[0])\n    print(\"sorted\")\n    sums = [item[1] for item in result]\n    \n    df_train = pd.concat(sums)\n    df_train = df_train.reset_index(drop=True)\n    #np.save() would be another option but this works for now\n    df_train.to_csv(\"my_train.csv.gz\", compression=\"gzip\")\n\n    df_train_scaled = pd.DataFrame(minmax.fit_transform(df_train))\n    df_train_scaled.to_csv(\"my_train_scaled.csv.gz\", compression=\"gzip\")\n    return df_train, df_train_scaled","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"46584de6b20af85588ed9d2cd0c9d50951c05e19"},"cell_type":"code","source":"ps = process_train()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"cd0caf26f0dddca9c95ff51fe3d95d6e748c91a2"},"cell_type":"markdown","source":"## Training dataset processing validation"},{"metadata":{"_uuid":"d8b4ee5cced5a62ea20a01f84cc30afb92683d8d"},"cell_type":"markdown","source":"And some brief look at the data itself, to see it is valid:"},{"metadata":{"trusted":true,"_uuid":"35cedcbce88c6895fcafc75a0fd436d6a824a327"},"cell_type":"code","source":"#first 10 rows of raw feature data\nps[0].head(10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"5bb76e10c8d0bb8c48a292168896a0da67d6fca7"},"cell_type":"code","source":"#same first 10 rows in scaled format\nps[1].head(10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"7c092a137dd69326c345f02159ea013ddccf32c8"},"cell_type":"code","source":"ps[1].values.shape","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"bbce2c46954635e15b269443a46ed6827bdf0218"},"cell_type":"markdown","source":"The above shows the shape of the generated data. In this case we have 22 features, so 22 columns. \n\n160 rows per signal (one \"bucket\" per row) so overall size matches:"},{"metadata":{"trusted":true,"_uuid":"e355487ea827711661231e1dc7ae36f6f060e6e3"},"cell_type":"code","source":"bkt_count*len(train_meta)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5cfe8d9bc961ce573982d95045881f790fb9ec1b"},"cell_type":"markdown","source":"The feature processing function in this kernel generates these features:\n* 5 for general bucket/bin statistics: bkt_mean, bkt_std, bkt_std_top, bkt_std_bottom, bkt_range\n* 3 for diff/lag in bucket/bin: bkt_diff_mean, bkt_diff_std, bkt_diff_sum\n* 7 percentiles: 0, 1, 25, 50, 75, 99, 100\n* 7 relative percentiles: 0, 1, 25, 50, 75, 99, 100\n-> total of 22 \"features\"\n\nTo show a bit how they all look together after scaled to range -1 to 1, a look at the first signal as processed into 160 buckets from 800k:"},{"metadata":{"trusted":true,"_uuid":"eeb680568e61d3b884206341623f0e89c2b16886"},"cell_type":"code","source":"ps[1][0:160].plot(figsize=(8,5))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"832817689b0d358d737e7a54a8b12f74991ba498"},"cell_type":"markdown","source":"All that is bit of a mess, so just the first bullet on its own -\n* 5 for general bucket/bin statistics: bkt_mean, bkt_std, bkt_std_top, bkt_std_bottom, bkt_range"},{"metadata":{"trusted":true,"_uuid":"841a7d2e4f3375391a3b3907f8f417d613ef35ea"},"cell_type":"code","source":"ps[1].iloc[:,0:5][0:160].plot()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ad3250b3c5999b57947ad3b9a83f317300794ec3"},"cell_type":"markdown","source":"Next the diff features:\n* 3 for diff/lag in bucket/bin: bkt_diff_mean, bkt_diff_std, bkt_diff_sum"},{"metadata":{"trusted":true,"_uuid":"557ef296b859c88ad5229bc22c9d723889debea4"},"cell_type":"code","source":"ps[1].iloc[:,5:8][0:160].plot()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"83c393b79858d0de90c6a0b5ff5639ab8ed2c19c"},"cell_type":"markdown","source":"Sum and average are overlapping, which is why only 2 lines show. So scaled average is the same as a scaled sum in this case. Should probably look at the othe features more closely as well, but that would be another story."},{"metadata":{"_uuid":"2dcc07eceadbe622ebde4c349acd6bf0a5d8e22d"},"cell_type":"markdown","source":"Next the percentiles:\n* 7 percentiles: 0, 1, 25, 50, 75, 99, 100"},{"metadata":{"trusted":true,"_uuid":"a0761907e1423395e308562472e3cd5250044f0b"},"cell_type":"code","source":"ps[1].iloc[:,8:15][0:160].plot()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d1cdf71f09f6861b6894617a966386a43e4d4455"},"cell_type":"markdown","source":"And the relative percentiles:\n* 7 relative percentiles: 0, 1, 25, 50, 75, 99, 100"},{"metadata":{"trusted":true,"_uuid":"51050b0296747a824cc0026e4b0a266345090e43"},"cell_type":"code","source":"ps[1].iloc[:,15:22][0:160].plot()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ad7fa2456c819cfe0109050280b99ba27111dd05"},"cell_type":"markdown","source":"## Combined dataset from above single-signal data"},{"metadata":{"_uuid":"4fbff0166db534838b802ecfcd9416e13896ab88"},"cell_type":"markdown","source":"The above showed an example for one signal with all the 22 features total.\n\nAs another dataset, I combine for each measurement id, the 3 signals into one row as features.\n\nThis allows running models where all 3 signals for a measirement id are treated as unified features, such as in the kernel I linked at the beginning.\n\n## Training-dataset combining:\n\nThe combine 3 signals code for the training dataset:"},{"metadata":{"trusted":true,"_uuid":"8d7c0e72bc0eedbe42550ab97e47525e7a439e46"},"cell_type":"code","source":"measurement_ids = train_meta[\"id_measurement\"].unique()\nrows = []\nfor mid in measurement_ids:\n    idx1 = mid*3\n    idx2 = idx1 + 1\n    idx3 = idx2 + 1\n    sig1_idx = idx1 * bkt_count\n    sig2_idx = idx2 * bkt_count\n    sig3_idx = idx3 * bkt_count\n    sig1_data = ps[1][sig1_idx:sig1_idx+bkt_count]\n    sig2_data = ps[1][sig2_idx:sig2_idx+bkt_count]\n    sig3_data = ps[1][sig3_idx:sig3_idx+bkt_count]\n    #this combines the above read 3*160 rows for 3 signals into 1 combined set with with 160 rows\n    #and from 22 features on 3*160 to 66 (=22*3) features on 160 rows.\n    row = np.concatenate([sig1_data, sig2_data, sig3_data], axis=1).flatten().reshape(bkt_count, sig1_data.shape[1]*3)\n    rows.append(row)\ndf_train_combined = pd.DataFrame(np.vstack(rows))\ndf_train_combined.to_csv(\"my_train_combined_scaled.csv.gz\", compression=\"gzip\")\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5d11de99feed099b0a91f36a114d45e2d936f565"},"cell_type":"markdown","source":"### Verification"},{"metadata":{"_uuid":"8bee593edeab783383fd50ceed7bf53dfd4bb4d5"},"cell_type":"markdown","source":"For verification, a look at the results to check the signal combination has worked:\n\nFirst the 2 first signals in the single signal dataset:"},{"metadata":{"trusted":true,"_uuid":"60f5a3f2ca94e25a13fb50c6272a1f27edaed74b"},"cell_type":"code","source":"#slot 1 (measurement 1, signal 1, rows 0-159) for single signal version\nps[1].iloc[:,15:22][0:160].plot()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e670421de0a76fef6863b4f429523033c5f19d70"},"cell_type":"code","source":"#slot 2 (measurement 1, signal 2, rows 160-319) for single signal version\nps[1].iloc[:,15:22][160:320].plot()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7eba15018489f7e6c13cb37655bab4b5febf43b1"},"cell_type":"markdown","source":"Now the same 2 signals in the combined dataset:"},{"metadata":{"trusted":true,"_uuid":"5b68e696c3688bd74b0651d782cf0e4a90ac6319"},"cell_type":"code","source":"#slot 1 (measurement 1, signal 1) for combined signal version\ndf_train_combined.iloc[:,15:22][0:160].plot()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"690f0d650bf889baa9652145290e26045a1d5529"},"cell_type":"code","source":"#slot 2 (measurement 1, signal 2) for combined signal version\ndf_train_combined.iloc[:,37:44][0:160].plot()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"68679d37d4355009536548acd6f3c3d7d5594f61"},"cell_type":"markdown","source":"One more, a look at the data contents to check also the values:"},{"metadata":{"trusted":true,"_uuid":"801a1503ff28d7dd22ac91a1decae59e6c2af18e"},"cell_type":"code","source":"#signal 1, single signal version\nps[1].iloc[0:4]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3cd234d3f7ac9d85abd9e23151f7c2471c49ba03"},"cell_type":"code","source":"#signal 2, single signal version\nps[1].iloc[160:164]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6604adf2a12dd848feef837645dd8d7a33d9aa9e"},"cell_type":"code","source":"#signal 3, single signal version\nps[1].iloc[320:324]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"2cee49df314057c55699c2ded7927e93f396f24d"},"cell_type":"code","source":"#signal 4 (or signal 1 for measurement id 2))\nps[1].iloc[480:484]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0e181f42bdf861a1be9f2d78e55cadc781d72288"},"cell_type":"markdown","source":"For comparison, the signals 1-3 for measurement id 1 in combined set:\n\nWith 22 features, signal 1 from above should be in columns 0-21, signal 2 in columns 22-43, and signal 3 in columns 44-65."},{"metadata":{"trusted":true,"_uuid":"30ef290bc87be8f696d36647576825016229c8d7"},"cell_type":"code","source":"df_train_combined.iloc[0:4]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f32119a8083929d127c9aca6fcb5415c4f211765"},"cell_type":"markdown","source":"And the second bucket at 160-320 rows should have signal 4 (or signal 1 for measurement id 2) from above, in columns 0-21:"},{"metadata":{"trusted":true,"_uuid":"a632e984c943bfd298a8e2281e7557d86a5a1f73"},"cell_type":"code","source":"df_train_combined.iloc[160:164]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"bc59fe346e9e5413fd2d028bd69ff34c2ec06cce"},"cell_type":"markdown","source":"All the above checks match, so I judge this as working as intended.\n\nSince this should now all be saved to disk, clean up some memory:"},{"metadata":{"trusted":true,"_uuid":"95c9ccd4085e2137bc147e67ef4c52c808e2fffc"},"cell_type":"code","source":"del ps\ndel df_train_combined","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f439c1783619ab41e913eb534ab9378b730b89d4"},"cell_type":"markdown","source":"## Test-dataset Processing and Features"},{"metadata":{"trusted":true,"_uuid":"dc1822b7648140ff28b561553ca6ead3d1a33585"},"cell_type":"markdown","source":"Now, the same multiprocessing elements for the test set as above for the training set:"},{"metadata":{"trusted":true,"_uuid":"26e4dcd6f53110cca6246faf148e53b32e0556c4"},"cell_type":"code","source":"def process_subtest(arg_tuple):\n    meta, idx = arg_tuple\n    df_sum = summarize_df_np(meta, \"test\", idx)\n    return idx, df_sum","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d85ad784853af760f1fa17c54056a394b3515aeb"},"cell_type":"code","source":"from multiprocessing import Pool\n\nnum_cores = 4\n\ndef process_test():\n    measurement_ids = test_meta[\"id_measurement\"].unique()\n    df_split = np.array_split(measurement_ids, num_cores)\n    chunk_size = len(df_split[0]) * 3\n    \n    chunk1 = create_chunk_indices(test_meta, 0, chunk_size)\n    chunk2 = create_chunk_indices(test_meta, 1, chunk_size)\n    chunk3 = create_chunk_indices(test_meta, 2, chunk_size)\n    chunk4 = create_chunk_indices(test_meta, 3, chunk_size)\n\n    all_chunks = [chunk1, chunk2, chunk3, chunk4]\n    \n    pool = Pool(num_cores)\n    result = pool.map(process_subtest, all_chunks)\n    result = sorted(result, key=lambda tup: tup[0])\n\n    sums = [item[1] for item in result]\n\n    df_test = pd.concat(sums)\n    df_test = df_test.reset_index(drop=True)\n    df_test.to_csv(\"my_test.csv.gz\", compression=\"gzip\")\n\n    df_test_scaled = pd.DataFrame(minmax.transform(df_test))\n    df_test_scaled.to_csv(\"my_test_scaled.csv.gz\", compression=\"gzip\")\n    return df_test, df_test_scaled","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"28d9f2689e8aba4d79538deeac22f9eb5a03c80d"},"cell_type":"code","source":"pst = process_test()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9df3bd57d810e5adeb251cb9c35bcbfcb1844ca9"},"cell_type":"code","source":"pst[0].head(10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"2be911bb8751de8b22758751716f83b2c81b9d9c"},"cell_type":"code","source":"pst[1].head(10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b1f8250127cf06398c39130c51b9fc62020f7911"},"cell_type":"code","source":"pst[1].values.shape","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9be3ea6b259bcd4804ce4bd187cabb0796352af5"},"cell_type":"markdown","source":"## A look at the processed test data"},{"metadata":{"_uuid":"54585cb398ffb62fadcb891af5b91edffe6644c2"},"cell_type":"markdown","source":"All-in-one figure for the mess:"},{"metadata":{"trusted":true,"_uuid":"f566dda562da3cea00aee9afb476612ecf98bd58"},"cell_type":"code","source":"pst[1][0:160].plot()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ee15bfbd0cd95c7c3a6f9d6292111a908ab4bcb0"},"cell_type":"markdown","source":"* The 5 general bucket/bin statistics:"},{"metadata":{"trusted":true,"_uuid":"97d816578cdb3d7f02b09dad60b602ce57af4408"},"cell_type":"code","source":"pst[1].iloc[:,0:5][0:160].plot()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1e3560389ef5881faaa32d2b5d818c6f24010e5e"},"cell_type":"markdown","source":"* The 3 for diff/lag in bucket/bin:"},{"metadata":{"trusted":true,"_uuid":"cf8a8a47d2d4a31394cded86ca228071cfde85b1"},"cell_type":"code","source":"pst[1].iloc[:,5:8][0:160].plot()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"49d8352f2d93a3b16d4740d90535a1c65110fd77"},"cell_type":"markdown","source":"* The 7 percentiles:"},{"metadata":{"trusted":true,"_uuid":"97a9d0ebc5c6ca550b3db6a826ee104c9f27b558"},"cell_type":"code","source":"pst[1].iloc[:,8:15][0:160].plot()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"053967684d451f36bdb713b0292fc6ed489a39d3"},"cell_type":"markdown","source":"* The 7 relative percentiles:"},{"metadata":{"trusted":true,"_uuid":"6f60865286ce0b99f57acd9f080bbb1806ffa242"},"cell_type":"code","source":"pst[1].iloc[:,15:22][0:160].plot()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"436820780e6fb3b4638327ef3e18a0b54b64c9cb"},"cell_type":"markdown","source":"## Test-dataset combinations"},{"metadata":{"_uuid":"f1743ee27fe038687ac99a4eb6f3a69c0c11ec81"},"cell_type":"markdown","source":"Similarly, combine 3 signals into one row as features for the combined version for test data:"},{"metadata":{"trusted":true,"_uuid":"e9ba68a1030c662ebb62a3c5ce34dd3c18aaf313"},"cell_type":"code","source":"measurement_ids = test_meta[\"id_measurement\"].unique()\nstart = measurement_ids[0]\nrows = []\nfor mid in measurement_ids:\n    #test measurement id's start from 2904 and indices at 0, so need to align\n    mid = mid - start\n    idx1 = mid*3\n    idx2 = idx1 + 1\n    idx3 = idx2 + 1\n    sig1_idx = idx1 * bkt_count\n    sig2_idx = idx2 * bkt_count\n    sig3_idx = idx3 * bkt_count\n    sig1_data = pst[1][sig1_idx:sig1_idx+bkt_count]\n    sig2_data = pst[1][sig2_idx:sig2_idx+bkt_count]\n    sig3_data = pst[1][sig3_idx:sig3_idx+bkt_count]\n    row = np.concatenate([sig1_data, sig2_data, sig3_data], axis=1).flatten().reshape(bkt_count, sig1_data.shape[1]*3)\n    rows.append(row)\ndf_test_combined = pd.DataFrame(np.vstack(rows))\ndf_test_combined.to_csv(\"my_test_combined_scaled.csv.gz\", compression=\"gzip\")\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a882e39ee5240f07a53385c5ceda16c95a47d4c7"},"cell_type":"markdown","source":"### A look at combined test-dataset"},{"metadata":{"_uuid":"fb15cab9cee8b17b73fc38255add3c7fc2f829c0"},"cell_type":"markdown","source":"Brief look at the combined test set signals to see it makes sense:"},{"metadata":{"trusted":true,"_uuid":"6ab1cb62553c53a0d13a17dcda59482f585e372b"},"cell_type":"code","source":"pst[1].iloc[0:4]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"cd69471efab36769dad7f5b5d06969904c295174"},"cell_type":"markdown","source":"vs."},{"metadata":{"trusted":true,"_uuid":"8a253c996db28274453d87c41fe63e719211f3c0"},"cell_type":"code","source":"df_test_combined.head(4)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"98a508bd54c9a9dade4f68b940b59bf08a83ab3b"},"cell_type":"markdown","source":"Check shape of combined signals to match number of unique measurements in test data:"},{"metadata":{"trusted":true,"_uuid":"eb854f7cb4711525d604b1907da8992810a8f5cb"},"cell_type":"code","source":"df_test_combined.shape","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4e8fceb21e885411d387469cbbbff575d3156750"},"cell_type":"markdown","source":"## Importing the Results to Other Kernels"},{"metadata":{"_uuid":"b8f7f71bcd2cf135529d6b11f89068351a9eff37"},"cell_type":"markdown","source":"To use the kernel output as a dataset in another (e.g., GPU) kernel, create a kernel and select \"+ Add Data\" on the right side panel."},{"metadata":{"trusted":true,"_uuid":"8d2916c442ca27283a2793a168996a2ff7d0ac75"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"cf1e27b87733abd350630c18c6ea7ea61ed8b65e"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}