{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport plotly.express as px","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-02-11T15:29:40.012853Z","iopub.execute_input":"2023-02-11T15:29:40.013287Z","iopub.status.idle":"2023-02-11T15:29:42.283563Z","shell.execute_reply.started":"2023-02-11T15:29:40.013254Z","shell.execute_reply":"2023-02-11T15:29:42.282290Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Preliminary EDA for the IceCube competition","metadata":{}},{"cell_type":"markdown","source":"The purpose of this notebook is to read, look and 'touch' the data available for IceCube Neutrino Competition. Before feature engineering and tinckering with ML models it is important to just check available data and make sense of it. This is what we're going to do in this notebook. ","metadata":{}},{"cell_type":"markdown","source":"## Contents:\n\n* Read available datasets\n* Batch Train Files\n* Sensor Geometry File\n* Test Meta File\n* Submission Sample File\n* Train Meta File\n* Overall Conclusion","metadata":{}},{"cell_type":"markdown","source":"### Read available datasets","metadata":{"execution":{"iopub.status.busy":"2023-02-10T22:49:24.473301Z","iopub.execute_input":"2023-02-10T22:49:24.473665Z","iopub.status.idle":"2023-02-10T22:49:24.477554Z","shell.execute_reply.started":"2023-02-10T22:49:24.473634Z","shell.execute_reply":"2023-02-10T22:49:24.476953Z"}}},{"cell_type":"markdown","source":"There is roughly 117 GB of available data broken down into a few hundred files. It's impractical for the scope of this project to read them all. What we're going to do is to read only one of each unique type of file to better inderstand what's inside.","metadata":{}},{"cell_type":"code","source":"train_parquet = pd.read_parquet('/kaggle/input/icecube-neutrinos-in-deep-ice/train/batch_1.parquet')\nsensor_geometry = pd.read_csv('/kaggle/input/icecube-neutrinos-in-deep-ice/sensor_geometry.csv')\ntest_meta = pd.read_parquet('/kaggle/input/icecube-neutrinos-in-deep-ice/test_meta.parquet')\ntrain_meta = pd.read_parquet('/kaggle/input/icecube-neutrinos-in-deep-ice/train_meta.parquet')\nsubmission_sample = pd.read_parquet('/kaggle/input/icecube-neutrinos-in-deep-ice/sample_submission.parquet')","metadata":{"execution":{"iopub.status.busy":"2023-02-11T15:29:42.285900Z","iopub.execute_input":"2023-02-11T15:29:42.286302Z","iopub.status.idle":"2023-02-11T15:30:30.368270Z","shell.execute_reply.started":"2023-02-11T15:29:42.286263Z","shell.execute_reply":"2023-02-11T15:30:30.367283Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Great! All files interesting to us were read and now we're ready to move to the EDA phase of the project. We're going to check every file separately and look if we can find anything of interest there.","metadata":{}},{"cell_type":"markdown","source":"### Batch Train Files","metadata":{}},{"cell_type":"markdown","source":"Batch train files are the corner stone of the whole competition.They contain information on the neutrino events detected by sensors. There are more than 600 individiual files in this category. We've read the first one - batch_1.parquet. Let's investigate what's under it's hood.","metadata":{}},{"cell_type":"code","source":"train_parquet.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-11T15:30:30.369716Z","iopub.execute_input":"2023-02-11T15:30:30.370044Z","iopub.status.idle":"2023-02-11T15:30:30.389325Z","shell.execute_reply.started":"2023-02-11T15:30:30.370013Z","shell.execute_reply":"2023-02-11T15:30:30.388227Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"What does it all mean? Kaggle metadata already contains information on what all these objects and attributes mean:\n\n* event_id (int): the event ID. Saved as the index column in parquet.\n* time (int): the time of the pulse in nanoseconds in the current event time window. The absolute time of a pulse has no relevance, and only the relative time with respect to other pulses within an event is of relevance.\n* sensor_id (int): the ID of which of the 5160 IceCube photomultiplier sensors recorded this pulse.\n* charge (float32): An estimate of the amount of light in the pulse, in units of photoelectrons (p.e.). A physical photon does not exactly result in a measurement of 1 p.e. but rather can take values spread around 1 p.e. As an example, a pulse with charge 2.7 p.e. could quite likely be the result of two or three photons hitting the photomultiplier tube around the same time. This data has float16 precision but is stored as float32 due to limitations of the version of pyarrow the data was prepared with.\n* auxiliary (bool): If True, the pulse was not fully digitized, is of lower quality, and was more likely to originate from noise. If False, then this pulse was contributed to the trigger decision and the pulse was fully digitized.","metadata":{}},{"cell_type":"markdown","source":"#### Datatypes, nulls, statistics","metadata":{}},{"cell_type":"code","source":"train_parquet.info()","metadata":{"execution":{"iopub.status.busy":"2023-02-11T15:30:30.390912Z","iopub.execute_input":"2023-02-11T15:30:30.392245Z","iopub.status.idle":"2023-02-11T15:30:30.411280Z","shell.execute_reply.started":"2023-02-11T15:30:30.392197Z","shell.execute_reply":"2023-02-11T15:30:30.409926Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"32 million rows in one table - not bad at all! Not a Titanic, for sure!","metadata":{}},{"cell_type":"code","source":"train_parquet.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2023-02-11T15:30:30.415731Z","iopub.execute_input":"2023-02-11T15:30:30.416431Z","iopub.status.idle":"2023-02-11T15:30:30.662728Z","shell.execute_reply.started":"2023-02-11T15:30:30.416389Z","shell.execute_reply":"2023-02-11T15:30:30.661471Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Well, the data was properly prepared for the competition. Hooray to the data engineering team!","metadata":{}},{"cell_type":"code","source":"train_parquet.describe()","metadata":{"execution":{"iopub.status.busy":"2023-02-11T15:30:30.664359Z","iopub.execute_input":"2023-02-11T15:30:30.664832Z","iopub.status.idle":"2023-02-11T15:30:34.820076Z","shell.execute_reply.started":"2023-02-11T15:30:30.664788Z","shell.execute_reply":"2023-02-11T15:30:34.818704Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We already can spot a few kinks in the distribution by a glimpse at the table above, but let's visualize the data to make a better sense of it.","metadata":{}},{"cell_type":"markdown","source":"#### Distribution histograms for each attribute","metadata":{}},{"cell_type":"code","source":"train_parquet.time.hist(bins=20, figsize=(7, 3));","metadata":{"execution":{"iopub.status.busy":"2023-02-11T15:30:34.821686Z","iopub.execute_input":"2023-02-11T15:30:34.822033Z","iopub.status.idle":"2023-02-11T15:30:35.844122Z","shell.execute_reply.started":"2023-02-11T15:30:34.822003Z","shell.execute_reply":"2023-02-11T15:30:35.842941Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Evidently, the majority of events lasts around 10k-15k nanoseconds, though there are quite a few outliers that are detected for much longer.","metadata":{}},{"cell_type":"code","source":"train_parquet.charge.hist(bins=50, figsize=(7, 3), range=(0,100));","metadata":{"execution":{"iopub.status.busy":"2023-02-11T15:30:35.845477Z","iopub.execute_input":"2023-02-11T15:30:35.845838Z","iopub.status.idle":"2023-02-11T15:30:37.067106Z","shell.execute_reply.started":"2023-02-11T15:30:35.845807Z","shell.execute_reply":"2023-02-11T15:30:37.065769Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Apparently, the distribution of charges is skewed towards zero. Or, in other words, the majority of detection events happened with a low energy particles.","metadata":{}},{"cell_type":"code","source":"train_parquet.auxiliary.value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-02-11T15:30:37.068614Z","iopub.execute_input":"2023-02-11T15:30:37.069065Z","iopub.status.idle":"2023-02-11T15:30:37.294800Z","shell.execute_reply.started":"2023-02-11T15:30:37.069030Z","shell.execute_reply":"2023-02-11T15:30:37.293853Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In this file, tha majority of objects were fully digitized (have False attribute) and that's a great news! ","metadata":{}},{"cell_type":"markdown","source":"#### Correlations","metadata":{}},{"cell_type":"code","source":"heatmap = train_parquet[['time', 'charge', 'sensor_id', 'auxiliary', 'sensor_id']].corr(method='spearman')\nsns.heatmap(data=heatmap, vmax=0.2, center=0, cmap='Spectral', square=True, linewidths=0.4, cbar_kws={\"shrink\": 0.5}, annot=True)\nplt.gcf().set_size_inches(10,10);\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-11T15:30:37.296326Z","iopub.execute_input":"2023-02-11T15:30:37.296952Z","iopub.status.idle":"2023-02-11T15:31:15.737021Z","shell.execute_reply.started":"2023-02-11T15:30:37.296917Z","shell.execute_reply":"2023-02-11T15:31:15.735776Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Finally, we determined that there is no strong correlations between attributes in the table. Though, there is some relation between the charge and the fact if it was fully digitised - which makes sense, since the stronger the signal, the better are the chances of a sensor to detect it.","metadata":{}},{"cell_type":"markdown","source":"#### Conclusion","metadata":{}},{"cell_type":"markdown","source":"We've read batch_1 file and investigated it's insides. \n\n* File is preprocessed and precleaned - at least on superficial level there is no need in additional preprocessing.\n\n* Distribution of numerical attributes isn't normal, at least for this file.\n\n* No strong correlation was found between attributes.\n\n* There is some correlation between the charge and it's detection.","metadata":{}},{"cell_type":"markdown","source":"### Sensor Geometry File","metadata":{}},{"cell_type":"markdown","source":"Let's move to the next file.","metadata":{}},{"cell_type":"code","source":"sensor_geometry.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-11T15:31:15.738802Z","iopub.execute_input":"2023-02-11T15:31:15.739612Z","iopub.status.idle":"2023-02-11T15:31:15.754956Z","shell.execute_reply.started":"2023-02-11T15:31:15.739565Z","shell.execute_reply":"2023-02-11T15:31:15.753917Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Just to be totally sure, let's check if there is any missing data.","metadata":{}},{"cell_type":"code","source":"sensor_geometry.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2023-02-11T15:31:15.756213Z","iopub.execute_input":"2023-02-11T15:31:15.756627Z","iopub.status.idle":"2023-02-11T15:31:15.775791Z","shell.execute_reply.started":"2023-02-11T15:31:15.756598Z","shell.execute_reply":"2023-02-11T15:31:15.774593Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Well, this is pretty straghtforward - the file consists of coordinates of sensors that detect the neutrino events.\n\nLet's have a look how they are positioned.","metadata":{}},{"cell_type":"code","source":"scatter3d = px.scatter_3d(sensor_geometry, x='x', y='y', z='z', opacity=0.7)\nscatter3d.update_traces(marker = dict(size = 2, symbol = \"diamond-open\"))\nscatter3d.update_coloraxes(showscale = False)\nscatter3d.update_layout(template = \"plotly\", font = dict(family = \"Arial\", size = 12, color = \"#9e97ff\"))\nscatter3d.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-11T15:31:15.777558Z","iopub.execute_input":"2023-02-11T15:31:15.778367Z","iopub.status.idle":"2023-02-11T15:31:16.935500Z","shell.execute_reply.started":"2023-02-11T15:31:15.778330Z","shell.execute_reply":"2023-02-11T15:31:16.934210Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Looks cool! And the chart does have similiarities with the picture in the competition description. File doesn't have anything else of interest for us, so let's move to the next one!","metadata":{}},{"cell_type":"markdown","source":"## Test Meta File","metadata":{}},{"cell_type":"code","source":"test_meta.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-11T15:31:16.940401Z","iopub.execute_input":"2023-02-11T15:31:16.940955Z","iopub.status.idle":"2023-02-11T15:31:16.953838Z","shell.execute_reply.started":"2023-02-11T15:31:16.940900Z","shell.execute_reply":"2023-02-11T15:31:16.952722Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Well, it's our test data. Apart from batch and id, it contains data on first/last row in the features dataframe belonging to this event.","metadata":{}},{"cell_type":"code","source":"test_meta.info()","metadata":{"execution":{"iopub.status.busy":"2023-02-11T15:31:16.955206Z","iopub.execute_input":"2023-02-11T15:31:16.955523Z","iopub.status.idle":"2023-02-11T15:31:16.971452Z","shell.execute_reply.started":"2023-02-11T15:31:16.955493Z","shell.execute_reply":"2023-02-11T15:31:16.969865Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The file is so small, because it is just an example of what our model is going to be tested against. So, let's move to the next file.","metadata":{}},{"cell_type":"markdown","source":"## Submission Sample File","metadata":{}},{"cell_type":"code","source":"submission_sample.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-11T15:31:16.973452Z","iopub.execute_input":"2023-02-11T15:31:16.973973Z","iopub.status.idle":"2023-02-11T15:31:16.987418Z","shell.execute_reply.started":"2023-02-11T15:31:16.973927Z","shell.execute_reply":"2023-02-11T15:31:16.986244Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So, this is the opposite of the previous file - submission_sample contains target data - azimuth and zenith - which we need to predict in the competition. What can be so difficult, right? Just insert 1s everywhere!\n\nLast, but not least - our train meta data.","metadata":{}},{"cell_type":"markdown","source":"## Train Meta File","metadata":{}},{"cell_type":"markdown","source":"It's one of the most important files we have (and the single biggest one as well). Let's have a look at what it has to offer.","metadata":{}},{"cell_type":"code","source":"train_meta.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-11T15:31:16.988736Z","iopub.execute_input":"2023-02-11T15:31:16.989165Z","iopub.status.idle":"2023-02-11T15:31:17.003515Z","shell.execute_reply.started":"2023-02-11T15:31:16.989110Z","shell.execute_reply":"2023-02-11T15:31:17.002537Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The data in the dataframe is pretty straightforward:\n* batch_id (int): the ID of the batch the event was placed into.\n* event_id (int): the event ID.\n* first/last_pulse_index (int): index of the first/last row in the features dataframe belonging to this event.\n* azimuth/zenith (float32): the azimuth/zenith angle in radians of the neutrino. A value between 0 and 2*pi for the azimuth and 0 and pi for zenith. The target columns. Not provided for the test set. The direction vector represented by zenith and azimuth points to where the neutrino came from.","metadata":{}},{"cell_type":"code","source":"train_meta.info()","metadata":{"execution":{"iopub.status.busy":"2023-02-11T16:09:40.084279Z","iopub.execute_input":"2023-02-11T16:09:40.085275Z","iopub.status.idle":"2023-02-11T16:09:40.100058Z","shell.execute_reply.started":"2023-02-11T16:09:40.085233Z","shell.execute_reply":"2023-02-11T16:09:40.098621Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The file is too big for Kaggle to handle properly, so before closer inspection let's take only a small subset of the data.","metadata":{}},{"cell_type":"code","source":"train_meta_small = train_meta.head(10000)","metadata":{"execution":{"iopub.status.busy":"2023-02-11T16:12:27.926213Z","iopub.execute_input":"2023-02-11T16:12:27.927231Z","iopub.status.idle":"2023-02-11T16:12:27.931443Z","shell.execute_reply.started":"2023-02-11T16:12:27.927190Z","shell.execute_reply":"2023-02-11T16:12:27.930629Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Correlations","metadata":{}},{"cell_type":"code","source":"heatmap = train_meta_small[['batch_id','event_id','first_pulse_index','last_pulse_index','azimuth','zenith']].corr(method='spearman')\nsns.heatmap(data=heatmap, vmax=0.2, center=0, cmap='Spectral', square=True, linewidths=0.4, cbar_kws={\"shrink\": 0.5}, annot=True)\nplt.gcf().set_size_inches(10,10);\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-11T16:16:00.015124Z","iopub.execute_input":"2023-02-11T16:16:00.015617Z","iopub.status.idle":"2023-02-11T16:16:00.432074Z","shell.execute_reply.started":"2023-02-11T16:16:00.015581Z","shell.execute_reply":"2023-02-11T16:16:00.430407Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Welp, time of beginning and end of the pulse are correlated. And they're correlated to the id - makes sense.","metadata":{}},{"cell_type":"markdown","source":"#### Distributions of attributes","metadata":{}},{"cell_type":"code","source":"fig, axs = plt.subplots(1, 4, figsize=(20, 3))\n\naxs[0].hist(train_meta_small[\"azimuth\"], bins=50)\naxs[0].set_title('Azimuth distribution')\n\naxs[1].hist(train_meta_small[\"zenith\"], bins=50)\naxs[1].set_title('Zenith distribution')\n\naxs[2].hist(train_meta_small[\"first_pulse_index\"], bins=50)\naxs[2].set_title('First Pulse distribution')\n\naxs[3].hist(train_meta_small[\"last_pulse_index\"], bins=50)\naxs[3].set_title('Last Pulse distribution')\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-11T16:26:46.701466Z","iopub.execute_input":"2023-02-11T16:26:46.702020Z","iopub.status.idle":"2023-02-11T16:26:47.540171Z","shell.execute_reply.started":"2023-02-11T16:26:46.701978Z","shell.execute_reply":"2023-02-11T16:26:47.538954Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Zenith distribution has a beautiful bell curve. Azimuths are pretty much evenly distributed. We can assume it has something to do with the geometry of the neutrino \"flight path\" - they can come from anywhere in terms of azimuth vectors, but there are some most probable zenith vectors which neutrinos are coming from (or rather detected).  \n\nTime of pulse is quite a mess.","metadata":{}},{"cell_type":"markdown","source":"#### Conclusion","metadata":{}},{"cell_type":"markdown","source":"We've checked the train_meta file and found that:\n\n* It is big (duh).\n* There is strong correlation between the start and end of the pulse and another - between time and event id.\n* There are some peculiarities in the geometry of the pulse that can be further investigated.","metadata":{}},{"cell_type":"markdown","source":"## Overall Conclusion","metadata":{}},{"cell_type":"markdown","source":"We've checked the available dataframes and found out a few interesting, though pretty surface level, correlations worth further investigation. The data itself is well preprocessed and doesn't require additional cleaning. \n\nSo, the important part is over - time for the fun, heh? ","metadata":{}}]}