{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport plotly.express as px","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-03-09T23:40:39.635386Z","iopub.execute_input":"2023-03-09T23:40:39.636115Z","iopub.status.idle":"2023-03-09T23:40:39.641678Z","shell.execute_reply.started":"2023-03-09T23:40:39.636068Z","shell.execute_reply":"2023-03-09T23:40:39.640553Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# IceCube Sensor Sensitivity: Feature Engineering","metadata":{}},{"cell_type":"markdown","source":"This project is based off [Paper Overview: Graph Neural Networks in IceCube notebook](https://www.kaggle.com/code/antonsevostianov/paper-overview-graph-neural-networks-in-icecube). We aren't going to discuss GNNs in this work, but concentrate on a few findings of the paper, which can, hopefully, help improve the models for the IceCube competition.\n\nThe purpose of this notebook is:\n\n1. Check how sensors of IceCube differ in performance.\n2. Create a **new feature** for future model - quantum efficiency (QE) aka sensor sensitivity/precision. \n\n### Contents:\n\n1. Brief IceCube Detectors Overview\n2. Breaking down detectors into groups\n3. Assigning numerical values of relative performance to the groups\n4. Conclusion","metadata":{}},{"cell_type":"markdown","source":"## IceCube Detectors\n\nAs the paper by R. Abbasi et al and the previously linked notebook discussed, IceCube Observatory isn't a monolith with homogenious sensors. All **detectors vary in their precision** and efficiency, therefore, this data can be used as a feature for future models.\n\nLet's have a look again at the diagram.\n\n<img src=\"https://i.postimg.cc/y8Cmr7Kg/IceCube.png\" width=\"400px\" height=\"500px\">\n\nMain takeaways are:\n\n* Ice around detectors aren't homogeneous - their absorption properties differs, depending on the depth. \n* [As a result] Different detectors have different performance/precision\n* *DeepCore enhanced quantum efficiency detectors* (marked green on the picture) are **the most precise**, according to the authors, they're at least 1.35 times more precise than standard detectors \n* Other *detectors beneath the Dust layer* [layer of ice with dust impurities] perform well, but worse than DeepCore ones\n* *Detectors above the Dust layer* perform worse than those beneath it\n* *Detectors inside the Dust layer* are **the worst**, precision wise","metadata":{}},{"cell_type":"markdown","source":"## Detector break down by groups\n\nAs described above, all sensors can be broken down into several groups by their relative performance.","metadata":{}},{"cell_type":"code","source":"sensor_geometry = pd.read_csv('/kaggle/input/icecube-neutrinos-in-deep-ice/sensor_geometry.csv')","metadata":{"execution":{"iopub.status.busy":"2023-03-09T23:40:39.658495Z","iopub.execute_input":"2023-03-09T23:40:39.659658Z","iopub.status.idle":"2023-03-09T23:40:39.672686Z","shell.execute_reply.started":"2023-03-09T23:40:39.659595Z","shell.execute_reply":"2023-03-09T23:40:39.671120Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Based on the information available, we can divide all sensors into following groups:\n\n* **DeepCore** sensors - the best quantum efficiency.\n* Detectors **under** the dust layer - second best QE\n* Detectors **above** the dust layer - third best QE\n* Detectors **inside** the dust layer - worst QE\n\nNow it's time to create  new feature, for now categorical.","metadata":{}},{"cell_type":"code","source":"def detector_group(x, z):\n    \"\"\"\n    Assigns values - deepcore, dustlayer, abovedust, underdust -\n    depending on sensor coordinates x and z\n    \"\"\"\n    \n    # Define functions for each sensor category\n    def is_deepcore(x, z):\n        return x in {57.2, -9.68, 31.25, 72.37, 113.19, 106.94, 41.6, -10.97} or \\\n               (x in {46.29, 194.34, 90.49, -32.96, -77.8, 1.71, 124.97} and \n                ((z <= 186.02 and z >= 95.91) or \n                 (z <= -157 and z >= -511)))\n    \n    def is_dustlayer(z):\n        return z <= 0 and z >= -155\n    \n    def is_abovedust(x, z):\n        return z > 0 and not is_deepcore(x, z) and not is_dustlayer(z)\n    \n    def is_underdust(x, z):\n        return not is_deepcore(x, z) and not is_dustlayer(z) and not is_abovedust(x, z)\n    \n    # Check which sensor category the coordinates belong to\n    if is_deepcore(x, z):\n        return \"deepcore\"\n    \n    if is_dustlayer(z):\n        return \"dustlayer\"\n    \n    if is_abovedust(x, z):\n        return \"abovedust\"\n    \n    return \"underdust\"\n","metadata":{"execution":{"iopub.status.busy":"2023-03-09T23:40:39.675169Z","iopub.execute_input":"2023-03-09T23:40:39.675521Z","iopub.status.idle":"2023-03-09T23:40:39.685879Z","shell.execute_reply.started":"2023-03-09T23:40:39.675486Z","shell.execute_reply":"2023-03-09T23:40:39.684710Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sensor_geometry['sensor_group'] = sensor_geometry.apply(lambda row: detector_group(row['x'], row['z']), axis=1)","metadata":{"execution":{"iopub.status.busy":"2023-03-09T23:40:39.687363Z","iopub.execute_input":"2023-03-09T23:40:39.687741Z","iopub.status.idle":"2023-03-09T23:40:39.803338Z","shell.execute_reply.started":"2023-03-09T23:40:39.687688Z","shell.execute_reply":"2023-03-09T23:40:39.801970Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sensor_geometry.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-09T23:40:39.806260Z","iopub.execute_input":"2023-03-09T23:40:39.806663Z","iopub.status.idle":"2023-03-09T23:40:39.822066Z","shell.execute_reply.started":"2023-03-09T23:40:39.806621Z","shell.execute_reply":"2023-03-09T23:40:39.820488Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"scatter3d = px.scatter_3d(sensor_geometry, x='x', y='y', z='z',color='sensor_group', opacity=0.7)\nscatter3d.update_traces(marker = dict(size = 2, symbol = \"diamond-open\"))\nscatter3d.update_coloraxes(showscale = False)\nscatter3d.update_layout(template = \"plotly\", font = dict(family = \"Arial\", size = 12, color = \"#9e97ff\"))\nscatter3d.show()","metadata":{"execution":{"iopub.status.busy":"2023-03-09T23:40:39.824160Z","iopub.execute_input":"2023-03-09T23:40:39.824738Z","iopub.status.idle":"2023-03-09T23:40:39.940096Z","shell.execute_reply.started":"2023-03-09T23:40:39.824663Z","shell.execute_reply":"2023-03-09T23:40:39.938645Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Et voilà! We've broken down all sensors into groups, depending on their quantum efficiency. Though, there is more to sensor QE, than just it's location within designated groups, even this rough division can give us additional information that can, probably, increase the score in the competition. ","metadata":{}},{"cell_type":"markdown","source":"### Detectors relative performance\n\nEach detector is unique and has slightly different performance comparing to others. Unfortunately there is little available information to assaign a certain value to every detector. However, albeit roughly, we still can give an estimate of relative performance of groups of sensors we created.\n\nAccording to R. Abbasi et al, DeepCore sensors are roughly 1.35 times more sensitive than \"normal\" detectors. We also know that dust layer detectors are roughly 0.5 - 0.6 time less sensitive. Based on this information we can engineer an additional feature - DOM's relative quantum efficiency.","metadata":{}},{"cell_type":"code","source":"def relative_qe(group):\n    \"\"\"\n    Returns a relative quantum efficiency of a sensor group according to the following rules:\n    - 1.35, if sensor is deepcore;\n    - 0.95, if sensor is abovedust;\n    - 1.05 if sensor is underdust;\n    - 0.6 if sensor is dustlayer\n    \"\"\"\n    if group == 'deepcore':\n        return 1.35\n    if group == 'abovedust':\n        return 0.95\n    if group == 'underdust':\n        return 1.05\n    else:\n        return 0.6","metadata":{"execution":{"iopub.status.busy":"2023-03-09T23:40:39.941661Z","iopub.execute_input":"2023-03-09T23:40:39.942371Z","iopub.status.idle":"2023-03-09T23:40:39.948694Z","shell.execute_reply.started":"2023-03-09T23:40:39.942332Z","shell.execute_reply":"2023-03-09T23:40:39.947186Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sensor_geometry['relative_qe'] = sensor_geometry['sensor_group'].apply(relative_qe)","metadata":{"execution":{"iopub.status.busy":"2023-03-09T23:40:39.950494Z","iopub.execute_input":"2023-03-09T23:40:39.950990Z","iopub.status.idle":"2023-03-09T23:40:39.962529Z","shell.execute_reply.started":"2023-03-09T23:40:39.950939Z","shell.execute_reply":"2023-03-09T23:40:39.961229Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sensor_geometry.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-09T23:40:39.966972Z","iopub.execute_input":"2023-03-09T23:40:39.967410Z","iopub.status.idle":"2023-03-09T23:40:39.987015Z","shell.execute_reply.started":"2023-03-09T23:40:39.967369Z","shell.execute_reply":"2023-03-09T23:40:39.985791Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Our dataset is ready! We can use one hot encoding to break down the categorical variable, if we want to, or just go with the numerical variable.","metadata":{}},{"cell_type":"markdown","source":"## Conclusion\n\nIn this notebook we were feature engineering a **sensor efficiency feature** that might come handy in training models for IceCube competiton. We have assigned a categorical and a numerical value to 4 different groups of detectors based on their relative performance.\n\nIf you're going to use this feature, please beware that numerical values are only rough estimates - so use your discretion while incorporating it into your models. While categorical column is probably fine, those numerical values might be hit or miss. However - this is the beauty of Kaggle competitions - you can try and see if it works in your model!\n\nPS. Please let me know if there are any mistakes. You most welcome to **share your experience - if this feature works or not for you**. ","metadata":{}}]}