{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\n# import numpy as np # linear algebra\n# import pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\n# import os\n# for dirname, _, filenames in os.walk('/kaggle/input'):\n#     for filename in filenames:\n#         print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-02-14T20:47:22.189326Z","iopub.execute_input":"2023-02-14T20:47:22.189847Z","iopub.status.idle":"2023-02-14T20:47:22.215201Z","shell.execute_reply.started":"2023-02-14T20:47:22.189736Z","shell.execute_reply":"2023-02-14T20:47:22.214185Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Files designations\n\n**[train/test]_meta.parquet**\n  \n**batch_id** (int): the ID of the batch the event was placed into.  \n**event_id** (int): the event ID.  \n**[first/last]_pulse_index** (int): index of the first/last row in the features dataframe belonging to this event.  \n**[azimuth/zenith]** (float32): the [azimuth/zenith] angle in radians of the neutrino. A value between 0 and 2*pi for the azimuth and 0 and pi for zenith. The target columns. Not provided for the test set. The direction vector represented by zenith and azimuth points to where the neutrino came from.  \n  \nNB: Other quantities regarding the event, such as the interaction point in x, y, z (vertex position), the neutrino energy, or the interaction type and kinematics are not included in the dataset.  \n  \n**[train/test]/batch_[n].parquet**   \n\nEach batch contains tens of thousands of events. Each event may contain thousands of pulses, each of which is the digitized output from a photomultiplier tube and occupies one row.\n\n**event_id** (int): the event ID. Saved as the index column in parquet.  \n**time** (int): the time of the pulse in nanoseconds in the current event time window. The absolute time of a pulse has no relevance, and only the relative time with respect to other pulses within an event is of relevance.  \n**sensor_id** (int): the ID of which of the 5160 IceCube photomultiplier sensors recorded this pulse.  \n**charge** (float32): An estimate of the amount of light in the pulse, in units of photoelectrons (p.e.). A physical photon does not exactly result in a measurement of 1 p.e. but rather can take values spread around 1 p.e. As an example, a pulse with charge 2.7 p.e. could quite likely be the result of two or three photons hitting the photomultiplier tube around the same time. This data has float16 precision but is stored as float32 due to limitations of the version of pyarrow the data was prepared with.  \n**auxiliary** (bool): If True, the pulse was not fully digitized, is of lower quality, and was more likely to originate from noise. If False, then this pulse was contributed to the trigger decision and the pulse was fully digitized.  \n  \n**sample_submission.parquet**   \n  \nAn example submission with the correct columns and properly ordered event IDs. The sample submission is provided in the parquet format so it can be read quickly but your final submission must be a csv.  \n  \n**sensor_geometry.csv**  \n  \nThe x, y, and z positions for each of the 5160 IceCube sensors. The row index corresponds to the sensor_idx feature of pulses. The x, y, and z coordinates are in units of meters, with the origin at the center of the IceCube detector. The coordinate system is right-handed, and the z-axis points upwards when standing at the South Pole. You can convert from these coordinates to azimuth and zenith with the following formulas (here the vector (x,y,z) is normalized):\n\nx = cos(azimuth) * sin(zenith)\ny = sin(azimuth) * sin(zenith)\nz = cos(zenith)","metadata":{}},{"cell_type":"markdown","source":"**Tasks and notes:**\n\n1) Need to combine azimuth/zenith angles (rad) of neutrinos and sensor_geometry data to get the full information about existing neutrinos' vectors (pulse location).  \n2) Sort impulses by time (time of the pulse in nanoseconds) - the final neutrino direction vector is constructed based on the time of occurrence of pulses  \n3) Neutrinos interact with each other through the weak nuclear force and influence each other's movement through gravitational forces. It might be intresting to study the influence of the `auxiliary`=True neutrinos on the direction of `auxiliary`=False neutrinos.  \n4) The motion of the neutrinos isn't directly related to the production of photoelectrons as the neutrinos don't need to be moving in a specific direction for the reaction of producing photoelectrons to occur.  \nThe direction of the neutrino's motion determines whether the photoelectrons are ejected in the same direction as the neutrino, or in the opposite direction. In the case of the opposite direction, the photoelectrons would be ejected in the opposite direction of the neutrino's motion.  ","metadata":{}},{"cell_type":"markdown","source":"# Import libraries","metadata":{}},{"cell_type":"code","source":"#! pip install pandas-profiling","metadata":{"execution":{"iopub.status.busy":"2023-02-14T20:47:22.217423Z","iopub.execute_input":"2023-02-14T20:47:22.218193Z","iopub.status.idle":"2023-02-14T20:47:22.224110Z","shell.execute_reply.started":"2023-02-14T20:47:22.218120Z","shell.execute_reply":"2023-02-14T20:47:22.222930Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport pandas_profiling\nfrom pandas_profiling import ProfileReport\nimport matplotlib.pyplot as plt\nimport plotly.express as px\nimport seaborn as sns\nimport matplotlib.image as mpimg\nimport numpy as np\nsns.set(style=\"darkgrid\")","metadata":{"execution":{"iopub.status.busy":"2023-02-14T20:47:22.226015Z","iopub.execute_input":"2023-02-14T20:47:22.226470Z","iopub.status.idle":"2023-02-14T20:47:26.496566Z","shell.execute_reply.started":"2023-02-14T20:47:22.226432Z","shell.execute_reply":"2023-02-14T20:47:26.495526Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data exploration - Batch 1","metadata":{"execution":{"iopub.status.busy":"2023-02-10T21:47:13.573462Z","iopub.execute_input":"2023-02-10T21:47:13.573881Z","iopub.status.idle":"2023-02-10T21:47:13.578999Z","shell.execute_reply.started":"2023-02-10T21:47:13.573848Z","shell.execute_reply":"2023-02-10T21:47:13.577812Z"}}},{"cell_type":"markdown","source":"## General information","metadata":{}},{"cell_type":"code","source":"train_meta_data = pd.read_parquet('/kaggle/input/icecube-neutrinos-in-deep-ice/train_meta.parquet')\ntrain_meta_data.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-14T20:47:26.498966Z","iopub.execute_input":"2023-02-14T20:47:26.499446Z","iopub.status.idle":"2023-02-14T20:48:09.854077Z","shell.execute_reply.started":"2023-02-14T20:47:26.499408Z","shell.execute_reply":"2023-02-14T20:48:09.852307Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_meta_data.info()","metadata":{"execution":{"iopub.status.busy":"2023-02-14T20:48:09.856581Z","iopub.execute_input":"2023-02-14T20:48:09.857684Z","iopub.status.idle":"2023-02-14T20:48:09.880667Z","shell.execute_reply.started":"2023-02-14T20:48:09.857632Z","shell.execute_reply":"2023-02-14T20:48:09.878983Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's check the data in train_batch_1 dataset:","metadata":{}},{"cell_type":"code","source":"train_batch_1_data = pd.read_parquet('/kaggle/input/icecube-neutrinos-in-deep-ice/train/batch_1.parquet')\ntrain_batch_1_data.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-14T20:48:09.882736Z","iopub.execute_input":"2023-02-14T20:48:09.883256Z","iopub.status.idle":"2023-02-14T20:48:13.288541Z","shell.execute_reply.started":"2023-02-14T20:48:09.883218Z","shell.execute_reply":"2023-02-14T20:48:13.287473Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_batch_1_data.info()","metadata":{"execution":{"iopub.status.busy":"2023-02-14T20:48:13.290611Z","iopub.execute_input":"2023-02-14T20:48:13.291276Z","iopub.status.idle":"2023-02-14T20:48:13.306137Z","shell.execute_reply.started":"2023-02-14T20:48:13.291236Z","shell.execute_reply":"2023-02-14T20:48:13.304672Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_batch_1_data.describe()","metadata":{"execution":{"iopub.status.busy":"2023-02-14T20:48:13.309289Z","iopub.execute_input":"2023-02-14T20:48:13.310459Z","iopub.status.idle":"2023-02-14T20:48:18.011410Z","shell.execute_reply.started":"2023-02-14T20:48:13.310313Z","shell.execute_reply":"2023-02-14T20:48:18.009456Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's do a quick exploratory analysis of the Batch_1:","metadata":{}},{"cell_type":"code","source":"profile = ProfileReport(train_batch_1_data)\nprofile","metadata":{"execution":{"iopub.status.busy":"2023-02-14T20:48:18.013860Z","iopub.execute_input":"2023-02-14T20:48:18.014538Z","iopub.status.idle":"2023-02-14T20:48:18.021524Z","shell.execute_reply.started":"2023-02-14T20:48:18.014492Z","shell.execute_reply":"2023-02-14T20:48:18.019256Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Overview:**  \n  \n1) Missing cells = 0  \n  \n2) Duplicate rows = 301 (need to delete)  \n  \n3) 200000 different events in one batch_1  \n  \n4) Some events are more common than others (are there more observations? more sensors pick up these events?)  \n    `event_id`=3266196 occurs 136 times  \n    `event_id`=3266175 occurs 124 times - this is two or more times more common than other events  \n  \n5) 5051 sensors in total (some sensors missed?)  \n  \n6) The histogram shows that sensors with a large `sensor_ID` value are more common. Especially sensors with id 4500-5000 and more. Perhaps a new, more sensitive type of sensor has been installed? Or the newer the sensor, the more events it detects (more sensitivity), but at the same time such a direct relationship isn’t visible up to 4000 sensor.\n  \n7) time - the relative time with respect to other pulses within an one event (nanoseconds).  \nBecause this is a relative value, most likely it makes no sense to look at the statistics for the entire batch. It’s necessary to look at the events separately.  \nBut it’s interesting to note that even the relative time (relative to other pulses) differs very much: min=5714 ns and a max=77785 ns (median=11815 ns). It’s possible that some of the sensors don’t work correctly enough and don’t capture all incoming pulses. With further analysis, you can see the dependence of the large relative time on the `sensor_id` - whether it is confirmed.  \n  \n8) `charge` - an estimate of the amount of light in the pulse, in units of photoelectrons (p.e.) - differs from min=0.025 to max=2762.025 p.e. in one pulse. It’s necessary to check whether such differences in the amount of light in the pulse are really possible, or are there more typical ones (will it be necessary to take into account the direction of the emission of photoelectrons when predicting the direction of neutrino motion? need to figure it out). Or is high values are sensors/ entries errors? (95-th percentile = 12.125)  \n  \n9) 28.2% of low quality pulses, 71.8% of fully digitilized pulses.  \n  \n10) `charge` weakly correlates with `auxiliary`, `time` weakly correlates with `sensor_id`.  \n\n    \n","metadata":{}},{"cell_type":"markdown","source":"## True and False pulses ratio ","metadata":{}},{"cell_type":"markdown","source":"Let's check what proportion of fully digitized pulses this batch batch_1 and `event_id`=24 contains:","metadata":{}},{"cell_type":"code","source":"train_batch_1_data = (train_batch_1_data\n                      .rename_axis('event_id')\n                      .reset_index())\n\ntrain_batch_1_data.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-14T20:48:18.030230Z","iopub.execute_input":"2023-02-14T20:48:18.030909Z","iopub.status.idle":"2023-02-14T20:48:18.646733Z","shell.execute_reply.started":"2023-02-14T20:48:18.030852Z","shell.execute_reply":"2023-02-14T20:48:18.645220Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"count_true = (train_batch_1_data[\n    (train_batch_1_data.event_id == 24) & \n    (train_batch_1_data.auxiliary == True)]).event_id.count()\n\ncount_false = (train_batch_1_data[\n    (train_batch_1_data.event_id == 24) & \n    (train_batch_1_data.auxiliary == False)]).event_id.count()\n\nprint(f'batch_1, event_id = 24, aux = True, count:'\\\n      f'{count_true}')\nprint(f'batch_1, event_id = 24, aux = False, count:'\\\n      f'{count_false}')","metadata":{"execution":{"iopub.status.busy":"2023-02-14T20:48:18.649253Z","iopub.execute_input":"2023-02-14T20:48:18.649798Z","iopub.status.idle":"2023-02-14T20:48:19.546502Z","shell.execute_reply.started":"2023-02-14T20:48:18.649745Z","shell.execute_reply":"2023-02-14T20:48:19.544984Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's check the ratio of `auxiliary`=True and `auxiliary`=False pulses for all events within batch 1:","metadata":{}},{"cell_type":"code","source":"batch_1_aux = (pd.pivot_table(train_batch_1_data, \n                             index='event_id', \n                             columns='auxiliary',\n                             aggfunc={'auxiliary':'count'})\n               .reset_index(level=0)\n              )\n\nbatch_1_aux.columns = batch_1_aux.columns.map(lambda x: x[1])\nbatch_1_aux = batch_1_aux.rename(columns={\n    '':'event_id', \n    False:'aux_false_count', \n    True:'aux_true_count'\n})\n\nbatch_1_aux['true_false_ratio'] = round(\n    batch_1_aux.aux_true_count / batch_1_aux.aux_false_count, \n    ndigits=4)\n\nbatch_1_aux","metadata":{"execution":{"iopub.status.busy":"2023-02-14T20:48:19.549279Z","iopub.execute_input":"2023-02-14T20:48:19.550864Z","iopub.status.idle":"2023-02-14T20:48:22.247870Z","shell.execute_reply.started":"2023-02-14T20:48:19.550566Z","shell.execute_reply":"2023-02-14T20:48:22.246450Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Draw a bar chart of the distribution of the low quality pulses (`auxiliary`=True) ratio relative to fully digitized pulses (`auxiliary`=False)","metadata":{}},{"cell_type":"code","source":"%matplotlib inline\nfig, ax = plt.subplots(figsize=(10, 6))\nax.scatter(x = batch_1_aux.aux_true_count, y = batch_1_aux.aux_false_count)\nplt.xlabel(\"Low quality pulses and noises qty\")\nplt.ylabel(\"Fully digitilized pulses qty\")\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-14T20:48:22.250110Z","iopub.execute_input":"2023-02-14T20:48:22.250556Z","iopub.status.idle":"2023-02-14T20:48:22.792475Z","shell.execute_reply.started":"2023-02-14T20:48:22.250509Z","shell.execute_reply":"2023-02-14T20:48:22.791104Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It can be seen that there're hundreds of times less noises than fully digitilized pulses. It's necessary to additionally check `events` where the number of `auxiliary`=False pulses is close to the number of `auxiliary`=True ones. And also see the distribution of incorrect pulses for specific sensors. Perhaps some of the sensors don't capture photoelectrons well enough?","metadata":{}},{"cell_type":"markdown","source":"**Next steps plan:**\n\n1) Check events where the number of `auxiliary`=False pulses is close to the number of `auxiliary`=True pulses  \n2) Check distribution of `auxiliary`=True pulses for specific `sensor_id`s  \n\n3) Check and delete duplicate rows in each batch  \n4) Check the hypothesis about newer and more sensitive sensors (ID 4000 and more), perhaps you can train the model on them and then correct the predictions for less sensitive ones","metadata":{}},{"cell_type":"markdown","source":"## Equal shares of true and false check","metadata":{}},{"cell_type":"markdown","source":"Let's check rows in batch_1 where the share of `auxiliary`=False pulses is equal to share of `axiliary`=True pulses (`true/false ratio` == 1)","metadata":{}},{"cell_type":"code","source":"batch_1_aux[batch_1_aux.true_false_ratio == 1]\n","metadata":{"execution":{"iopub.status.busy":"2023-02-14T21:06:06.218844Z","iopub.execute_input":"2023-02-14T21:06:06.219357Z","iopub.status.idle":"2023-02-14T21:06:06.242936Z","shell.execute_reply.started":"2023-02-14T21:06:06.219314Z","shell.execute_reply":"2023-02-14T21:06:06.241616Z"},"trusted":true},"execution_count":null,"outputs":[]}]}