{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"This Data preprocessing goal\n\n* Creating train dataset by combining meta_data, sensor data and batch_data\n* and I use numpy for it\n\n\n\nthe pulse was not fully digitized, is of lower quality, and was more likely to originate from noise. At Kamiokande, the analog signals from the photodetectors are digitized and the data is read out to a computer. In other words, signals (pulses) that have not been digitized are not used in calculations. Therefore, it is not considered necessary for this calculation. (If there is enough power, digitizing the unprocessed pulses will greatly increase the amount of useful data!)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1132983%2F6891ec67d9d40315637b1b292c3a486b%2FExample_event.png?generation=1666631264548536&alt=media)","metadata":{}},{"cell_type":"code","source":"%matplotlib inline\n\nimport os\nimport glob\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport time\n\nPATH_DATASET = \"/kaggle/input/icecube-neutrinos-in-deep-ice/\"\ncolumns = []\n\neventid_charge_cols = []\neventid_time_cols = []\nfor i in range(5160):\n    eventid_charge_cols.append(\n    'sensor_id' + '_' + str(i) + '_TC')\n    eventid_charge_cols.append(\n    'sensor_id' + '_' + str(i) + '_TM')\ncolumns.extend(eventid_charge_cols)\ncolumns.extend(eventid_time_cols)\ncolumns.extend(['event_id', 'azimuth', 'zenith'])\n\ndel eventid_charge_cols, eventid_time_cols\nprint(len(columns))","metadata":{"execution":{"iopub.status.busy":"2023-01-22T09:11:24.412151Z","iopub.execute_input":"2023-01-22T09:11:24.412626Z","iopub.status.idle":"2023-01-22T09:11:24.456020Z","shell.execute_reply.started":"2023-01-22T09:11:24.412497Z","shell.execute_reply":"2023-01-22T09:11:24.454746Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Preparation/preproccessing","metadata":{}},{"cell_type":"markdown","source":"[train/test]_meta.parquet\n\n* batch_id (int): the ID of the batch the event was placed into.\n* event_id (int): the event ID.\n* [first/last]_pulse_index (int): index of the first/last row in the features dataframe belonging to this event.\n* [azimuth/zenith] (float32): the [azimuth/zenith] angle in radians of the neutrino. A value between 0 and 2*pi for the azimuth and 0 and pi for zenith. The target columns. Not provided for the test set. The direction vector represented by zenith and azimuth points to where the neutrino came from.\n","metadata":{}},{"cell_type":"markdown","source":"[train/test]/batch_[n].parquet Each batch contains tens of thousands of events. Each event may contain thousands of pulses, each of which is the digitized output from a photomultiplier tube and occupies one row.\n\n* event_id (int): the event ID. Saved as the index column in parquet.\n* time (int): the time of the pulse in nanoseconds in the current event time window. The absolute time of a pulse has no relevance, and only the relative time with respect to other pulses within an event is of relevance.\n* sensor_id (int): the ID of which of the 5160 IceCube photomultiplier sensors recorded this pulse.\n* charge (float32): An estimate of the amount of light in the pulse, in units of photoelectrons (p.e.). A physical photon does not exactly result in a measurement of 1 p.e. but rather can take values spread around 1 p.e. As an example, a pulse with charge 2.7 p.e. could quite likely be the result of two or three photons hitting the photomultiplier tube around the same time. This data has float16 precision but is stored as float32 due to limitations of the version of pyarrow the data was prepared with.\n* auxiliary (bool): If True, the pulse was not fully digitized, is of lower quality, and was more likely to originate from noise. If False, then this pulse was contributed to the trigger decision and the pulse was fully digitized.","metadata":{}},{"cell_type":"code","source":"meta_train = pd.read_parquet(os.path.join(PATH_DATASET, \"train_meta.parquet\"))","metadata":{"execution":{"iopub.status.busy":"2023-01-22T09:11:24.458411Z","iopub.execute_input":"2023-01-22T09:11:24.458906Z","iopub.status.idle":"2023-01-22T09:12:06.434942Z","shell.execute_reply.started":"2023-01-22T09:11:24.458862Z","shell.execute_reply":"2023-01-22T09:12:06.433669Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"event_list = meta_train['event_id'][:10]","metadata":{"execution":{"iopub.status.busy":"2023-01-22T09:12:06.436398Z","iopub.execute_input":"2023-01-22T09:12:06.436748Z","iopub.status.idle":"2023-01-22T09:12:06.445048Z","shell.execute_reply.started":"2023-01-22T09:12:06.436717Z","shell.execute_reply":"2023-01-22T09:12:06.443756Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"trn_batch = pd.read_parquet(os.path.join(PATH_DATASET, 'train/batch_1.parquet'))","metadata":{"execution":{"iopub.status.busy":"2023-01-22T09:12:39.185187Z","iopub.execute_input":"2023-01-22T09:12:39.185659Z","iopub.status.idle":"2023-01-22T09:12:42.153693Z","shell.execute_reply.started":"2023-01-22T09:12:39.185613Z","shell.execute_reply":"2023-01-22T09:12:42.152612Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def preprocessing_TrainData(event_id: int, trn_batch: pd.DataFrame, meta_train: pd.DataFrame):\n    start = time.time()\n    \n    meta_train_eventid = meta_train[meta_train['event_id'] == event_id]\n    del meta_train\n    df_all_eventid = trn_batch.merge(meta_train_eventid, on = 'event_id', how = 'left')\n    del meta_train_eventid\n    \n    # the total charge variables for each sensor_id\n    event_charge = np.zeros((5160))\n    sensor_ar = np.array(df_all_eventid['sensor_id'], dtype=int)\n    charge_ar = np.array(df_all_eventid['charge'], dtype=float)\n    ar_eventid_charge = np.bincount(sensor_ar, charge_ar)[1:]\n    event_charge[:sensor_ar.max()] = ar_eventid_charge\n\n    # Now create the first time variables for each sensor_id\n    event_time = np.zeros((5161))\n    sensor = np.array(df_all_eventid.sort_values(by='time')['sensor_id'], dtype=int)\n    sample_time = df_all_eventid['time'].to_numpy()\n    _, indices = np.unique(sensor, return_index=True)\n    u_time = sample_time[indices]\n    sensor_idx = np.unique(sensor)\n    \n    event_time[sensor_idx] = u_time\n    event_time = event_time[1:]\n    \n    az, ze = df_all_eventid['azimuth'].max(), df_all_eventid['zenith'].max()\n    del df_all_eventid\n\n    # combine all tables together\n    variables_all = np.hstack((event_time, event_charge,\n                               np.array(event_id), np.array(az), np.array(ze)))                          \n    end = time.time()\n    \n    print(f'Elapsed time is {end - start} seconds')\n#     print(variables_all.shape)\n    return variables_all.reshape(1, -1)","metadata":{"execution":{"iopub.status.busy":"2023-01-22T09:24:16.237076Z","iopub.execute_input":"2023-01-22T09:24:16.237547Z","iopub.status.idle":"2023-01-22T09:24:16.250678Z","shell.execute_reply.started":"2023-01-22T09:24:16.237496Z","shell.execute_reply":"2023-01-22T09:24:16.249229Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"event_list","metadata":{"execution":{"iopub.status.busy":"2023-01-22T09:24:17.380894Z","iopub.execute_input":"2023-01-22T09:24:17.381301Z","iopub.status.idle":"2023-01-22T09:24:17.388539Z","shell.execute_reply.started":"2023-01-22T09:24:17.381266Z","shell.execute_reply":"2023-01-22T09:24:17.387595Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"numpy_X = np.ndarray((10, len(columns)))\nfor i, id in enumerate(event_list):\n    numpy_X[i] = preprocessing_TrainData(event_id = id, trn_batch = trn_batch, meta_train = meta_train)","metadata":{"execution":{"iopub.status.busy":"2023-01-22T09:25:32.420367Z","iopub.execute_input":"2023-01-22T09:25:32.420820Z","iopub.status.idle":"2023-01-22T09:29:56.737954Z","shell.execute_reply.started":"2023-01-22T09:25:32.420785Z","shell.execute_reply":"2023-01-22T09:29:56.736613Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X = pd.DataFrame(numpy_X, columns = columns)\n\noutput_dir = \"/kaggle/working/\"\nX.to_csv(os.path.join(output_dir, \"sample_preprocessed_train.csv\"))","metadata":{"execution":{"iopub.status.busy":"2023-01-22T09:29:57.763931Z","iopub.execute_input":"2023-01-22T09:29:57.764292Z","iopub.status.idle":"2023-01-22T09:29:57.924263Z","shell.execute_reply.started":"2023-01-22T09:29:57.764260Z","shell.execute_reply":"2023-01-22T09:29:57.922914Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"event_id length has 131953924. In this preprocess, It will take about many many days to complete all tasks !!!!\n\nIf you have any other suggestions for improvement, we would appreciate it if you could let us know.","metadata":{}}]}