{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Info\n\nThe purpose of this notebook is to provide a very basic intro to this competition.\n\nThe target of this Kaggle competition is to predict the direction from which a neutrino was detected by the IceCube neutrino observatory. The direction is represented by two angles: azimuth and zenith. These angles are not provided in the test set, and are the **target columns**.\n\nThe given features are the event details and pulse information. The event details include the:\n- batch ID, \n- event ID, \n- first pulse index, and last pulse index of each event. \n\nThe pulse information include:\n- time, \n- sensor ID, \n- charge, \n- auxiliary, and sensor geometry for each pulse. \n\nThe sensor geometry information includes the x, y, and z positions of each of the 5160 IceCube sensors. These positions are used to convert from the coordinate system to azimuth and zenith angles.","metadata":{}},{"cell_type":"markdown","source":"# Insights\n\nHere is the collection of insights collected through the notebook:\n\n- For a given `azimuth` and `zenith`, there is a single combination of sensor_ids. This mean that it might be a good idea to convert the sensor_id to `sensor_azimuth` and `sensor_zenith` variables which tell you the what is the azimuth and zenith of a given sensor position.\n- Features are time-ordered, which means that sequence direction is probably important to the model. Time (position) embedding might be a good idea here.\n- Majority of charges are below the value of 4, altough values up to 2700 occur. These outliers might have to be treated.\n- There seem to be a difference in a distribution between train and test data in charge. Majority of test data is for charges lower than 2, and the largest test charge is 4.78. Maybe it's a good idea to train only on data for charge < 4.78. Or you can clamp the charge from 0 to 5.\n- Charge occurrs in discrete values, with a delta Q of 0.05.\n- It seems that there is a majority of `auxiliary=True` events by a factor of 2.\n- It's very important to remove the center of mass from the points to correctly estimate the angles!\n- The simple regression of z vs xy yields accurate predictions of one of the angles, given that the other is known. This might indicate that a 3-stage training can be efficient: in first stage, you explicitly give the model the known azimuth (or zenith) and train the model with frozen zenith (azimuth) weights. In the second stage, you freeze azimuth (zenith) weights and train zenith (azimuth) weights. In the final stage, you train the model with no frozen weights.\n- Majority of the data has less than 256 context length, which can be a good input for a model.\n- Majority of the `auxiliary`=False data has less than 128 context length! This might mean that it might be efficient to take only non-auxiliary data for the input\n","metadata":{}},{"cell_type":"code","source":"# import dependencies\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport plotly.graph_objs as go\nfrom sklearn.linear_model import LinearRegression\n","metadata":{"execution":{"iopub.status.busy":"2023-02-27T21:42:38.962976Z","iopub.execute_input":"2023-02-27T21:42:38.963472Z","iopub.status.idle":"2023-02-27T21:42:40.433361Z","shell.execute_reply.started":"2023-02-27T21:42:38.963361Z","shell.execute_reply":"2023-02-27T21:42:40.431903Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Explore the features\n\nLet's first explore the features that we are dealt with.\n\nLet's load the train_meta.parquet","metadata":{}},{"cell_type":"code","source":"# load /kaggle/input/icecube-neutrinos-in-deep-ice/train_meta.parquet\ndf = pd.read_parquet(\"/kaggle/input/icecube-neutrinos-in-deep-ice/train_meta.parquet\")\ndf","metadata":{"execution":{"iopub.status.busy":"2023-02-27T21:42:40.435418Z","iopub.execute_input":"2023-02-27T21:42:40.436092Z","iopub.status.idle":"2023-02-27T21:43:20.062373Z","shell.execute_reply.started":"2023-02-27T21:42:40.436042Z","shell.execute_reply":"2023-02-27T21:43:20.060897Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# load /kaggle/input/icecube-neutrinos-in-deep-ice/test_meta.parquet\ndf_test = pd.read_parquet(\"/kaggle/input/icecube-neutrinos-in-deep-ice/test_meta.parquet\")\ndf_test[\"batch_id\"].nunique()","metadata":{"execution":{"iopub.status.busy":"2023-02-27T21:43:20.064108Z","iopub.execute_input":"2023-02-27T21:43:20.064550Z","iopub.status.idle":"2023-02-27T21:43:20.087662Z","shell.execute_reply.started":"2023-02-27T21:43:20.064512Z","shell.execute_reply":"2023-02-27T21:43:20.086508Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\nOkay, so the `event_id` is the id of an event (obviously haha), and `batch_id` is the id of a batch where the data is contained. This means that for `batch_id=1` training data features are in `train/batch_1.parquet`. Let's check how this looks like:","metadata":{}},{"cell_type":"code","source":"df_example = pd.read_parquet(\"/kaggle/input/icecube-neutrinos-in-deep-ice/train/batch_1.parquet\").reset_index()\ndf_example[df_example[\"event_id\"]==24]","metadata":{"execution":{"iopub.status.busy":"2023-02-27T21:43:20.090505Z","iopub.execute_input":"2023-02-27T21:43:20.091885Z","iopub.status.idle":"2023-02-27T21:43:23.985759Z","shell.execute_reply.started":"2023-02-27T21:43:20.091837Z","shell.execute_reply":"2023-02-27T21:43:23.984770Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_example_test = pd.read_parquet(\"/kaggle/input/icecube-neutrinos-in-deep-ice/train/batch_1.parquet\").reset_index()\ndf_example_test\n# df_example_test[df_example_test[\"event_id\"]==2092]","metadata":{"execution":{"iopub.status.busy":"2023-02-27T21:43:23.987400Z","iopub.execute_input":"2023-02-27T21:43:23.988067Z","iopub.status.idle":"2023-02-27T21:43:25.096978Z","shell.execute_reply.started":"2023-02-27T21:43:23.988027Z","shell.execute_reply":"2023-02-27T21:43:25.095648Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Great, so what information does this event tell us? \n\n- `sensor_id`: tells us which of the 5160 IceCube sensors was triggered by this event. We'll explore this in more depth\n- `time`: this tells the \"time of the pulse in nanoseconds in the current event time window\". I will assume that this depends on the velocity of the neutrino \n- `charge`: an estimate of the charge. This is a float number\n- `auxiliary`: This means that this event might actually occur from the noise","metadata":{}},{"cell_type":"markdown","source":"Let's first explore the `sensor_id`","metadata":{}},{"cell_type":"code","source":"df_sensor_geometry = pd.read_csv(\"/kaggle/input/icecube-neutrinos-in-deep-ice/sensor_geometry.csv\")\ndf_sensor_geometry\n","metadata":{"execution":{"iopub.status.busy":"2023-02-27T21:43:25.098381Z","iopub.execute_input":"2023-02-27T21:43:25.099653Z","iopub.status.idle":"2023-02-27T21:43:25.127873Z","shell.execute_reply.started":"2023-02-27T21:43:25.099554Z","shell.execute_reply":"2023-02-27T21:43:25.126832Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So, it's just an identifier for the grid point. \nHow many points in each direction we have?","metadata":{}},{"cell_type":"code","source":"# Get the unique values of x, y and z\nx_unique = df_sensor_geometry[\"x\"].unique()\ny_unique = df_sensor_geometry[\"y\"].unique()\nz_unique = df_sensor_geometry[\"z\"].unique()\nprint(\"shapes\", x_unique.shape, y_unique.shape, z_unique.shape)\n# print min and max values of x, y and z\nprint(\"min\", x_unique.min(), y_unique.min(), z_unique.min())\nprint(\"max\", x_unique.max(), y_unique.max(), z_unique.max())","metadata":{"execution":{"iopub.status.busy":"2023-02-27T21:43:25.129394Z","iopub.execute_input":"2023-02-27T21:43:25.129819Z","iopub.status.idle":"2023-02-27T21:43:25.139412Z","shell.execute_reply.started":"2023-02-27T21:43:25.129783Z","shell.execute_reply":"2023-02-27T21:43:25.136836Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The geometry is symmetric around origin.\n\n\nThis might mean that it's useful to work in spherical coordinates. Note about that: the formulas on this link (https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/data) are not 100% corrent: they should read\n\n```latex\nx = r * cos(azimuth) * sin(zenith)\ny = r * sin(azimuth) * sin(zenith)\nz = r * cos(zenith)\n```\nwhere $r=\\sqrt{x^2 + y^2 + z^2}$, and not \n\n```\nx = cos(azimuth) * sin(zenith)\ny = sin(azimuth) * sin(zenith)\nz = cos(zenith)\n```\n\nThe inverse relation, going from $(x, y, z)$ to $(r, \\theta, \\varphi)$ can be obtained as follows:\n\n$r=\\sqrt{x^2 + y^2 + z^2}$,\n\n$\\theta = \\text{arctan}(\\sqrt{x^2 + y^2} / z)$\n\n$\\varphi = \\text{arctan}(y / x)$\n\nSome great visualization is provided here 😊: https://keisan.casio.com/exec/system/1359533867 \n\n","metadata":{}},{"cell_type":"markdown","source":"## insight 💡\n\nFor a given `azimuth` and `zenith`, there is a single combination of sensor_ids. This mean that it might be a good idea to convert the sensor_id to `sensor_azimuth` and `sensor_zenith` variables which tell you the what is the azimuth and zenith of a given sensor position.","metadata":{}},{"cell_type":"markdown","source":"## Further exploration\n\n","metadata":{}},{"cell_type":"code","source":"df_tmp = pd.merge(df_example, df_sensor_geometry, on=\"sensor_id\")\ndf_tmp_all_events = pd.merge(df_tmp, df[df[\"batch_id\"]==1].reset_index()[[\"event_id\", \"azimuth\", \"zenith\"]], on=\"event_id\")\n# df_tmp_all_events = df_tmp[df_tmp[\"auxiliary\"]==False].reset_index()\ndf_tmp_all_events[\"event_id\"].unique()","metadata":{"execution":{"iopub.status.busy":"2023-02-27T21:43:25.141475Z","iopub.execute_input":"2023-02-27T21:43:25.142351Z","iopub.status.idle":"2023-02-27T21:43:44.985201Z","shell.execute_reply.started":"2023-02-27T21:43:25.142227Z","shell.execute_reply":"2023-02-27T21:43:44.982324Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tmp = df_tmp_all_events[df_tmp_all_events[\"event_id\"] == 1128590].reset_index()\ntmp","metadata":{"execution":{"iopub.status.busy":"2023-02-27T21:43:44.988755Z","iopub.execute_input":"2023-02-27T21:43:44.989253Z","iopub.status.idle":"2023-02-27T21:43:45.912322Z","shell.execute_reply.started":"2023-02-27T21:43:44.989207Z","shell.execute_reply":"2023-02-27T21:43:45.910662Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot(df_event):\n    \n    x = np.array(df_event[\"x\"] - np.average(df_event[\"x\"]))\n    y = np.array(df_event[\"y\"] - np.average(df_event[\"y\"]))\n    z = np.array(df_event[\"z\"] - np.average(df_event[\"z\"]))\n    \n    r = np.sqrt(x**2 + y**2 + z**2)\n    \n    charge = df_event[\"charge\"].values\n    \n    # Create a linear regression object and fit the data\n    reg = LinearRegression().fit(np.column_stack((x, y)), z, sample_weight=1/r)\n\n    # Extract the coefficients of the plane\n    m_x, m_y = reg.coef_  # slopes\n    b = reg.intercept_  # y-intercept\n    \n    # Create a scatter3d trace with x, y, and z positions from the df_event DataFrame\n    trace_points = go.Scatter3d(\n        x=x,\n        y=y,\n        z=z,\n        mode='markers',\n        marker=dict(\n            # size=2,\n            size=3,# np.clip(df_event['charge'], 0, 5),\n            # width=df_event['charge'],\n            color='blue',\n            colorscale='Viridis',\n            opacity=0.8\n        ),\n        name='Points'\n    )\n    \n    \n\n    # Convert the azimuth and zenith angles to radians\n    \n    phi = df_event['azimuth']\n    theta =   df_event['zenith']\n\n    # Calculate the line endpoints in x, y, and z coordinates\n    x2 = np.linspace(-800, 800, len(df_event)) * (np.cos(phi) * np.sin(theta))\n    y2 = np.linspace(-800, 800, len(df_event)) * (np.sin(theta) * np.sin(phi))\n    z2 = np.linspace(-800, 800, len(df_event)) * np.cos(theta)\n    \n    \n    trace_fitten_points = go.Scatter3d(\n        x=x2,\n        y=y2,\n        z=m_x * x2 + m_y * y2 + b,\n        mode='lines',\n        line=dict(\n            color='purple',\n            width=10\n        ),\n        name='Fitted regression line'\n    )\n    \n    # Create a line trace with the calculated endpoints\n    trace_line = go.Scatter3d(\n        x=x2,\n        y=y2,\n        z=z2,\n        mode='lines',\n        line=dict(\n            color='red',\n            width=10\n        ),\n        name='Target Line'\n    )\n    \n    \n    # Combine the scatter and line traces into a np.single plot\n    fig = go.Figure(data=[trace_points, trace_line, trace_fitten_points])\n\n    # Add axis labels and a title\n    fig.update_layout(\n        scene=dict(\n            xaxis_title='X',\n            yaxis_title='Y',\n            zaxis_title='Z'\n        ),\n        title='Interactive 3D plot'\n    )\n\n    # Display the plot in Jupyter Notebook or in a web browser\n    fig.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-02-27T21:43:45.916726Z","iopub.execute_input":"2023-02-27T21:43:45.917104Z","iopub.status.idle":"2023-02-27T21:43:45.935376Z","shell.execute_reply.started":"2023-02-27T21:43:45.917077Z","shell.execute_reply":"2023-02-27T21:43:45.932973Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_event = df_tmp_all_events[df_tmp_all_events[\"event_id\"] == df_tmp_all_events[\"event_id\"].unique()[0]].reset_index(drop=True)\nplot(df_event)","metadata":{"execution":{"iopub.status.busy":"2023-02-27T21:43:45.937425Z","iopub.execute_input":"2023-02-27T21:43:45.937887Z","iopub.status.idle":"2023-02-27T21:43:46.423857Z","shell.execute_reply.started":"2023-02-27T21:43:45.937845Z","shell.execute_reply":"2023-02-27T21:43:46.422079Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_event = df_tmp_all_events[df_tmp_all_events[\"event_id\"] == df_tmp_all_events[\"event_id\"].unique()[1]].reset_index(drop=True)\nplot(df_event)","metadata":{"execution":{"iopub.status.busy":"2023-02-27T21:43:46.425482Z","iopub.execute_input":"2023-02-27T21:43:46.425869Z","iopub.status.idle":"2023-02-27T21:43:46.639345Z","shell.execute_reply.started":"2023-02-27T21:43:46.425832Z","shell.execute_reply":"2023-02-27T21:43:46.637575Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_event = df_tmp_all_events[df_tmp_all_events[\"event_id\"] == df_tmp_all_events[\"event_id\"].unique()[2]].reset_index(drop=True)\nplot(df_event)","metadata":{"execution":{"iopub.status.busy":"2023-02-27T21:43:46.641276Z","iopub.execute_input":"2023-02-27T21:43:46.641961Z","iopub.status.idle":"2023-02-27T21:43:46.854860Z","shell.execute_reply.started":"2023-02-27T21:43:46.641914Z","shell.execute_reply":"2023-02-27T21:43:46.853505Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_event = df_tmp_all_events[df_tmp_all_events[\"event_id\"] == df_tmp_all_events[\"event_id\"].unique()[4]].reset_index(drop=True)\nplot(df_event)","metadata":{"execution":{"iopub.status.busy":"2023-02-27T21:43:46.856417Z","iopub.execute_input":"2023-02-27T21:43:46.856779Z","iopub.status.idle":"2023-02-27T21:43:47.062834Z","shell.execute_reply.started":"2023-02-27T21:43:46.856745Z","shell.execute_reply":"2023-02-27T21:43:47.062101Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_event = df_tmp_all_events[df_tmp_all_events[\"event_id\"] == 1128590].reset_index(drop=True)\nplot(df_event)","metadata":{"execution":{"iopub.status.busy":"2023-02-27T21:43:47.063889Z","iopub.execute_input":"2023-02-27T21:43:47.064322Z","iopub.status.idle":"2023-02-27T21:43:47.130859Z","shell.execute_reply.started":"2023-02-27T21:43:47.064296Z","shell.execute_reply":"2023-02-27T21:43:47.126824Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_event = df_tmp_all_events[df_tmp_all_events[\"event_id\"] == 1799490].reset_index(drop=True)\nplot(df_event)","metadata":{"execution":{"iopub.status.busy":"2023-02-27T21:43:47.132059Z","iopub.execute_input":"2023-02-27T21:43:47.132410Z","iopub.status.idle":"2023-02-27T21:43:47.183939Z","shell.execute_reply.started":"2023-02-27T21:43:47.132378Z","shell.execute_reply":"2023-02-27T21:43:47.182977Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# insight 💡💡💡\n\n\n- It's very important to remove the center of mass from the points to correctly estimate the angles!\n- The simple regression of z vs xy yields accurate predictions of one of the angles, given that the other is known. This might indicate that a 3-stage training can be efficient: in first stage, you explicitly give the model the known azimuth (or zenith) and train the model with frozen zenith (azimuth) weights. In the second stage, you freeze azimuth (zenith) weights and train zenith (azimuth) weights. In the final stage, you train the model with no frozen weights.","metadata":{}},{"cell_type":"markdown","source":"## Explore the `time`\nLet's explore the `time` variable in the example event.\n\nHere is the result of a `.describe()` and histogram.","metadata":{}},{"cell_type":"code","source":"\ndf_example[\"time\"].describe()","metadata":{"execution":{"iopub.status.busy":"2023-02-27T21:43:47.185525Z","iopub.execute_input":"2023-02-27T21:43:47.186135Z","iopub.status.idle":"2023-02-27T21:43:47.695020Z","shell.execute_reply.started":"2023-02-27T21:43:47.186098Z","shell.execute_reply":"2023-02-27T21:43:47.693693Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.hist(df_example[\"time\"], bins=20, alpha=0.5, range=(5000, 30000), density=True, label='train')\nplt.hist(df_example_test[\"time\"], bins=20, alpha=0.5, range=(5000, 30000), density=True, label='test');\nplt.xlim(5000, 30000)\nplt.legend(loc=\"best\")\nplt.xlabel(\"time\")","metadata":{"execution":{"iopub.status.busy":"2023-02-27T21:43:47.696578Z","iopub.execute_input":"2023-02-27T21:43:47.697014Z","iopub.status.idle":"2023-02-27T21:43:48.798641Z","shell.execute_reply.started":"2023-02-27T21:43:47.696977Z","shell.execute_reply":"2023-02-27T21:43:48.797014Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"OK, but does the time ordering matter?","metadata":{}},{"cell_type":"code","source":"x = df_example[df_example[\"event_id\"]==24][\"time\"].values\nplt.plot(x, label=24)\n\nx = df_example[df_example[\"event_id\"]==72][\"time\"].values\nplt.plot(x, label=72)\n\nx = df_example[df_example[\"event_id\"]==67][\"time\"].values\nplt.plot(x, label=67)\n\nx = df_example[df_example[\"event_id\"]==41][\"time\"].values\nplt.plot(x, label=41)\n\nx = df_example[df_example[\"event_id\"]==2147483597][\"time\"].values\nplt.plot(x, label=2147483597)\n\nplt.ylabel(\"cumulative time\")\nplt.xlabel(\"index\")\n","metadata":{"execution":{"iopub.status.busy":"2023-02-27T21:43:48.803428Z","iopub.execute_input":"2023-02-27T21:43:48.805009Z","iopub.status.idle":"2023-02-27T21:43:49.081050Z","shell.execute_reply.started":"2023-02-27T21:43:48.804947Z","shell.execute_reply":"2023-02-27T21:43:49.079347Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## insight 💡\n\nFeatures are time-ordered, which means that sequence direction is probably important to the model. Time (position) embedding might be a good idea here\n\nLet's find out what is the maximum time to get more insight into ","metadata":{}},{"cell_type":"code","source":"max_times = df_example.groupby(\"event_id\").max(\"time\")[\"time\"]\nplt.hist(max_times.values, bins=100);\nprint(np.max(max_times))\nplt.xlabel(\"time\")","metadata":{"execution":{"iopub.status.busy":"2023-02-27T21:43:49.082888Z","iopub.execute_input":"2023-02-27T21:43:49.083253Z","iopub.status.idle":"2023-02-27T21:43:51.555919Z","shell.execute_reply.started":"2023-02-27T21:43:49.083219Z","shell.execute_reply":"2023-02-27T21:43:51.554056Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Explore the `charge`\nLet's explore the `charge` variable in the example event.\n\nHere is the result of a `.describe()` and histogram.","metadata":{}},{"cell_type":"code","source":"df_example[\"charge\"].describe()","metadata":{"execution":{"iopub.status.busy":"2023-02-27T21:43:51.558315Z","iopub.execute_input":"2023-02-27T21:43:51.558878Z","iopub.status.idle":"2023-02-27T21:43:52.677658Z","shell.execute_reply.started":"2023-02-27T21:43:51.558834Z","shell.execute_reply":"2023-02-27T21:43:52.675921Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Oh, there seem to be some outliers in the charge (no pun intended 🤣)\n\n\n\n## insight 💡\n\nMajority of charges are below the value of 4, altough values up to 2700 occur. These outliers might have to be treated.\n\n\n\nMaybe a log10 of the charge might be a good feature.","metadata":{}},{"cell_type":"code","source":"plt.hist(df_example[\"charge\"], alpha=0.5, bins=100, density=True, range=(0, 10), label=\"Train\");\nplt.hist(df_example_test[\"charge\"],alpha=0.5,  bins=100,density=True, range=(0, 10), label=\"Test\");\n\n\nplt.xlabel(\"charge\")\nplt.legend(loc=\"best\")\n","metadata":{"execution":{"iopub.status.busy":"2023-02-27T21:43:52.679996Z","iopub.execute_input":"2023-02-27T21:43:52.680485Z","iopub.status.idle":"2023-02-27T21:43:53.992224Z","shell.execute_reply.started":"2023-02-27T21:43:52.680445Z","shell.execute_reply":"2023-02-27T21:43:53.989862Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"percentage_gt_2_train = np.sum(df_example[\"charge\"] > 2) / len(df_example[\"charge\"])\npercentage_gt_2_test = np.sum(df_example_test[\"charge\"] > 2) / len(df_example_test[\"charge\"])\npercentage_gt_2_train, percentage_gt_2_test","metadata":{"execution":{"iopub.status.busy":"2023-02-27T21:43:53.994897Z","iopub.execute_input":"2023-02-27T21:43:53.995650Z","iopub.status.idle":"2023-02-27T21:43:54.082753Z","shell.execute_reply.started":"2023-02-27T21:43:53.995605Z","shell.execute_reply":"2023-02-27T21:43:54.081676Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"np.max(df_example_test[\"charge\"])","metadata":{"execution":{"iopub.status.busy":"2023-02-27T21:43:54.084266Z","iopub.execute_input":"2023-02-27T21:43:54.084678Z","iopub.status.idle":"2023-02-27T21:43:54.163185Z","shell.execute_reply.started":"2023-02-27T21:43:54.084642Z","shell.execute_reply":"2023-02-27T21:43:54.161970Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"np.sum(df_example[\"charge\"] > 4.78) / len(df_example[\"charge\"])","metadata":{"execution":{"iopub.status.busy":"2023-02-27T21:43:54.164682Z","iopub.execute_input":"2023-02-27T21:43:54.165093Z","iopub.status.idle":"2023-02-27T21:43:54.207855Z","shell.execute_reply.started":"2023-02-27T21:43:54.165003Z","shell.execute_reply":"2023-02-27T21:43:54.205848Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## insight 💡\n\nThere seem to be a difference in a distribution between train and test data in charge. Majority of test data is for charges lower than 2, and the largest test charge is 4.78. Maybe it's a good idea to train only on data for charge < 4.78.","metadata":{}},{"cell_type":"code","source":"plt.hist(np.log10(df_example[\"charge\"]), bins=100, range=(0, 4));\nplt.xlabel(\"log10(charge)\");","metadata":{"execution":{"iopub.status.busy":"2023-02-27T21:43:54.209488Z","iopub.execute_input":"2023-02-27T21:43:54.209949Z","iopub.status.idle":"2023-02-27T21:43:55.235129Z","shell.execute_reply.started":"2023-02-27T21:43:54.209911Z","shell.execute_reply":"2023-02-27T21:43:55.233957Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.hist(np.log10(df_example[\"charge\"]), bins=100, range=(0, 4));\nplt.yscale('log')\nplt.xlabel(\"log10(charge)\");","metadata":{"execution":{"iopub.status.busy":"2023-02-27T21:43:55.237949Z","iopub.execute_input":"2023-02-27T21:43:55.238545Z","iopub.status.idle":"2023-02-27T21:43:57.041462Z","shell.execute_reply.started":"2023-02-27T21:43:55.238497Z","shell.execute_reply":"2023-02-27T21:43:57.039604Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's see if the charge is discretized:","metadata":{}},{"cell_type":"code","source":"unique_charges = df_example[\"charge\"].unique()\nprint(f\"there are {len(set(unique_charges))} unique charges\")\ncharges = np.sort(unique_charges)\ndx = np.diff(charges)\n# round to 3 decimal places\ndx = np.round(dx, 3)\n# find unique values\ndx = np.unique(dx)\n\nprint(\"First 10 delta Q:\"), dx[:10]","metadata":{"execution":{"iopub.status.busy":"2023-02-27T21:43:57.048625Z","iopub.execute_input":"2023-02-27T21:43:57.049069Z","iopub.status.idle":"2023-02-27T21:43:57.311874Z","shell.execute_reply.started":"2023-02-27T21:43:57.049039Z","shell.execute_reply":"2023-02-27T21:43:57.310580Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## insight 💡\n\nCharge occurrs in discrete values, with a delta Q of 0.05.\n\n","metadata":{}},{"cell_type":"markdown","source":"# Exploring `auxiliary`\n\n","metadata":{}},{"cell_type":"code","source":"# count number of True and False elements in df_example[\"auxiliary\"]\ndf_tmp = pd.read_parquet(\"/kaggle/input/icecube-neutrinos-in-deep-ice/train/batch_1.parquet\").reset_index()\nprint(df_tmp[\"auxiliary\"].value_counts())\n\ndf_tmp = pd.read_parquet(\"/kaggle/input/icecube-neutrinos-in-deep-ice/train/batch_2.parquet\").reset_index()\nprint(df_tmp[\"auxiliary\"].value_counts())\n\ndf_tmp = pd.read_parquet(\"/kaggle/input/icecube-neutrinos-in-deep-ice/train/batch_3.parquet\").reset_index()\nprint(df_tmp[\"auxiliary\"].value_counts())","metadata":{"execution":{"iopub.status.busy":"2023-02-27T21:43:57.314075Z","iopub.execute_input":"2023-02-27T21:43:57.314554Z","iopub.status.idle":"2023-02-27T21:44:06.826502Z","shell.execute_reply.started":"2023-02-27T21:43:57.314513Z","shell.execute_reply":"2023-02-27T21:44:06.824293Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n\nIt seems that there is a majority of `auxiliary=True` events by a factor of 2.","metadata":{}},{"cell_type":"markdown","source":"# Understanding the targets\n\nA line in 3D space in this competition is represented as an intersection of `azimuth` and `zenith`. In maths, these are the angles $\\varphi$ and $\\theta$ of the spherical system:\n\n\n\n ![https://mathworld.wolfram.com/images/eps-svg/SphericalCoordinates_1201.svg](https://mathworld.wolfram.com/images/eps-svg/SphericalCoordinates_1201.svg) .\n \n \n So, a fixed `azimuth` is a cone, and a fixed `zenith` angle is a half-plane. If you intersect the two, you get a line. You can check the middle and right plots below, which are defined as constant zenith and constant azimuth respectively: \n \n \n ![https://math.libretexts.org/@api/deki/files/3550/CNX_Calc_Figure_12_07_019.jpg?revision=1&size=bestfit&width=623&height=262](https://math.libretexts.org/@api/deki/files/3550/CNX_Calc_Figure_12_07_019.jpg?revision=1&size=bestfit&width=623&height=262)\n ","metadata":{}},{"cell_type":"markdown","source":"# Exploring `azimuth`","metadata":{}},{"cell_type":"code","source":"plt.hist(df['azimuth'].values / (np.pi), bins=100);\nplt.xlabel(\"azimuth (or phi, in units of pi)\")","metadata":{"execution":{"iopub.status.busy":"2023-02-27T21:44:06.828613Z","iopub.execute_input":"2023-02-27T21:44:06.829101Z","iopub.status.idle":"2023-02-27T21:44:09.615720Z","shell.execute_reply.started":"2023-02-27T21:44:06.829056Z","shell.execute_reply":"2023-02-27T21:44:09.613233Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Exploring `zenith`","metadata":{}},{"cell_type":"code","source":"plt.hist(df['zenith'].values / (np.pi), bins=100);\nplt.xlabel(\"zenith (or theta, in units of pi)\")","metadata":{"execution":{"iopub.status.busy":"2023-02-27T21:44:09.618124Z","iopub.execute_input":"2023-02-27T21:44:09.618707Z","iopub.status.idle":"2023-02-27T21:44:12.174512Z","shell.execute_reply.started":"2023-02-27T21:44:09.618657Z","shell.execute_reply":"2023-02-27T21:44:12.173144Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Meta-feature exploration\n\nLet's explore how large is the \"context\" of a given target signal.","metadata":{}},{"cell_type":"code","source":"contextual_lengths = df[\"last_pulse_index\"] - df[\"first_pulse_index\"]\nprint(\"min/max contextual length:\", np.min(contextual_lengths), np.max(contextual_lengths))","metadata":{"execution":{"iopub.status.busy":"2023-02-27T21:44:12.176876Z","iopub.execute_input":"2023-02-27T21:44:12.177343Z","iopub.status.idle":"2023-02-27T21:44:13.077419Z","shell.execute_reply.started":"2023-02-27T21:44:12.177302Z","shell.execute_reply":"2023-02-27T21:44:13.075367Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.hist(contextual_lengths, bins=50, range=(0, 256));\nplt.xlabel(\"contextual length\")","metadata":{"execution":{"iopub.status.busy":"2023-02-27T21:44:13.079848Z","iopub.execute_input":"2023-02-27T21:44:13.080329Z","iopub.status.idle":"2023-02-27T21:44:15.072650Z","shell.execute_reply.started":"2023-02-27T21:44:13.080277Z","shell.execute_reply":"2023-02-27T21:44:15.070883Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f\"number of elements greater than 256 context size: {np.sum(contextual_lengths > 256) / len(contextual_lengths):.3f}\")","metadata":{"execution":{"iopub.status.busy":"2023-02-27T21:44:15.074774Z","iopub.execute_input":"2023-02-27T21:44:15.075176Z","iopub.status.idle":"2023-02-27T21:44:15.220000Z","shell.execute_reply.started":"2023-02-27T21:44:15.075138Z","shell.execute_reply":"2023-02-27T21:44:15.217743Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## insight 💡 \nMajority of the data has less than 256 context length, which can be a good input for a model.","metadata":{}},{"cell_type":"code","source":"# Make the same exploration, but taking only `auxiliary`=False\ndf_batch = pd.read_parquet(\"/kaggle/input/icecube-neutrinos-in-deep-ice/train/batch_1.parquet\").reset_index()\ndf_batch = df_batch[df_batch[\"auxiliary\"]==False].reset_index()\ncontextual_lengths = df_batch.groupby('event_id').size().values\nprint(\"min/max contextual length:\", np.min(contextual_lengths), np.max(contextual_lengths))\n\nplt.hist(contextual_lengths, bins=50, range=(0, 256));\nplt.xlabel(\"contextual size\")\nplt.xlim(0, 128)\n\nprint(f\"number of elements greater than 128 context size: {np.sum(contextual_lengths > 128) / len(contextual_lengths):.3f}\")","metadata":{"execution":{"iopub.status.busy":"2023-02-27T21:44:15.222817Z","iopub.execute_input":"2023-02-27T21:44:15.223227Z","iopub.status.idle":"2023-02-27T21:44:18.133934Z","shell.execute_reply.started":"2023-02-27T21:44:15.223190Z","shell.execute_reply":"2023-02-27T21:44:18.131920Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## insight 💡 \nMajority of the `auxiliary`=False data has less than 128 context length! This might mean that it might be efficient to take only non-auxiliary data for the input","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}