{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<h1 style=\"text-align: center;\"><b>EDA Simplified:<span style=\"color:#6f94b0\n;\"> IceCube Neutrinos in Deep Ice</span></b></h1>","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #6f94b0; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Introduction</h2>\n\nAround the universe, we see particles everywhere. Photons, Electrons, Neutrons, and you name it, Neutrinos, because they are one of the most abundant particles around our world. It's similar like the electrons but they didn't have electrical charge and mostly zero mass, as it posed a problem for detecting them. And as for collecting information about them, scientists must roughly calculate the direction of neutrino time events. However, the research based on the events from the neutrino directions sparked more problems as the existing solutions aren't perfect until the far future when the neutrino particles are fast but either accurate or inaccurate depending on the computational coasts. But there's a catch. The IceCube Neutrino Observatory from the University of Wisconsin-Madison is one of the first detectors to detect the direction of neutrinos deep into the bottom from the icy layer in Antarctica, and hosted a competition to reconstruct the time events of the Neutrino movement, but with that info gathered toward us, let's drill into our casual EDA!","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #6f94b0; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Imports & Setup</h2>\n\nWhen we get started on our Data Analysis on the Neutrino directions, we import the pandas module as pd and the numpy module as np casually for data science and linear algebra. Next for plotting the data out, we load the plotly module followed by the express submodule as px and the altair module as alt.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\n\nimport plotly.express as px\nimport altair as alt","metadata":{"execution":{"iopub.status.busy":"2023-02-11T00:22:45.591703Z","iopub.execute_input":"2023-02-11T00:22:45.592212Z","iopub.status.idle":"2023-02-11T00:22:47.182390Z","shell.execute_reply.started":"2023-02-11T00:22:45.592105Z","shell.execute_reply":"2023-02-11T00:22:47.181090Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Once we complete importing the four necessary modules, we characterize the sensor_geometry_df dataframe to the pd module's read_csv function, reading out the file directory leading to the sensor_geometry.csv file from the competition data. However, we noticed that there's mostly parquet files in the competition, so we characterize the train_df to use the read_parquet function from the pd module to read out the train_meta.parquet file from the competition data. With that, we display out the first five rows of the two dataframes we created, with the head function.","metadata":{}},{"cell_type":"code","source":"sensor_geometry_df = pd.read_csv(\"/kaggle/input/icecube-neutrinos-in-deep-ice/sensor_geometry.csv\")\ntrain_df = pd.read_parquet(\"/kaggle/input/icecube-neutrinos-in-deep-ice/train_meta.parquet\")","metadata":{"execution":{"iopub.status.busy":"2023-02-11T00:22:47.184530Z","iopub.execute_input":"2023-02-11T00:22:47.184857Z","iopub.status.idle":"2023-02-11T00:23:37.077564Z","shell.execute_reply.started":"2023-02-11T00:22:47.184832Z","shell.execute_reply":"2023-02-11T00:23:37.073586Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sensor_geometry_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-11T00:23:37.083169Z","iopub.execute_input":"2023-02-11T00:23:37.083728Z","iopub.status.idle":"2023-02-11T00:23:37.156843Z","shell.execute_reply.started":"2023-02-11T00:23:37.083682Z","shell.execute_reply":"2023-02-11T00:23:37.148745Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-11T00:23:37.163858Z","iopub.execute_input":"2023-02-11T00:23:37.165641Z","iopub.status.idle":"2023-02-11T00:23:37.215625Z","shell.execute_reply.started":"2023-02-11T00:23:37.165565Z","shell.execute_reply":"2023-02-11T00:23:37.207639Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3 style=\"background-color: #6f94b0; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 8px; border-radius: 15px;\">Basic Dataframe Analysis</h3>\n\nAfter we finished creating the two dataframes, let's get started on visualizing the basic information about them! First, let's find the number of data counts with the len function to the sensor_geometry_df and the train_df dataframes.","metadata":{}},{"cell_type":"code","source":"len(sensor_geometry_df)","metadata":{"execution":{"iopub.status.busy":"2023-02-11T00:23:37.228185Z","iopub.execute_input":"2023-02-11T00:23:37.230777Z","iopub.status.idle":"2023-02-11T00:23:37.276250Z","shell.execute_reply.started":"2023-02-11T00:23:37.230741Z","shell.execute_reply":"2023-02-11T00:23:37.267850Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(train_df)","metadata":{"execution":{"iopub.status.busy":"2023-02-11T00:23:37.278997Z","iopub.execute_input":"2023-02-11T00:23:37.282469Z","iopub.status.idle":"2023-02-11T00:23:37.331743Z","shell.execute_reply.started":"2023-02-11T00:23:37.282402Z","shell.execute_reply":"2023-02-11T00:23:37.327385Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see, there are 5160 data entities listed in the sensor_geometry_df dataframe and not only that, we saw a whooping 131953924 data entities calculated in the train_df dataframe. That's because there's a lot of neutrino movement gathered by scientists which they must be reconstructed because of inaccurate information.","metadata":{}},{"cell_type":"markdown","source":"Now let's visualize and calculate the number of NaN values in the two dataframes! We use each of the two dataframes to find the NaN values with the isna function, followed by adding them up with the sum function.","metadata":{}},{"cell_type":"code","source":"sensor_geometry_df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2023-02-11T00:23:37.340427Z","iopub.execute_input":"2023-02-11T00:23:37.340972Z","iopub.status.idle":"2023-02-11T00:23:37.375559Z","shell.execute_reply.started":"2023-02-11T00:23:37.340929Z","shell.execute_reply":"2023-02-11T00:23:37.370167Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2023-02-11T00:23:37.379066Z","iopub.execute_input":"2023-02-11T00:23:37.380866Z","iopub.status.idle":"2023-02-11T00:23:38.756617Z","shell.execute_reply.started":"2023-02-11T00:23:37.380824Z","shell.execute_reply":"2023-02-11T00:23:38.752163Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Perfect! There are no NaN values present in both of the two dataframes. In other words, the data gathered in the two dataframes shows us that the data from the IceCube Neutrino Laboratory is sent to the data warehouses in University of Wisconsin-Madison.","metadata":{}},{"cell_type":"markdown","source":"Finally, let's find the number of columns in the two dataframes! We use the shape attribute to the two dataframes to find the shape of them along inserting the slice index of 1 to return the last shape.","metadata":{}},{"cell_type":"code","source":"sensor_geometry_df.shape[1]","metadata":{"execution":{"iopub.status.busy":"2023-02-11T00:23:38.758534Z","iopub.execute_input":"2023-02-11T00:23:38.758885Z","iopub.status.idle":"2023-02-11T00:23:38.773599Z","shell.execute_reply.started":"2023-02-11T00:23:38.758854Z","shell.execute_reply":"2023-02-11T00:23:38.771652Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.shape[1]","metadata":{"execution":{"iopub.status.busy":"2023-02-11T00:23:38.785703Z","iopub.execute_input":"2023-02-11T00:23:38.790482Z","iopub.status.idle":"2023-02-11T00:23:38.834923Z","shell.execute_reply.started":"2023-02-11T00:23:38.790418Z","shell.execute_reply":"2023-02-11T00:23:38.830642Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the outputs of the two cells above, we found out that there are 4 columns in the sensor_geometry_df dataframe and 6 columns in the train_df dataframe. In other words, the small number of columns specified in the two dataframes hinted us that they are finding the relevant data of the neutrino movements based on position or pulses.","metadata":{}},{"cell_type":"markdown","source":"And just like that, we're finished with the basic dataframe analysis of the sensor_geometry_df and the train_df dataframes! Without further ado, let's head onwards to visualize the two dataframes we created first, and then we visualize the batches specified in the competition data!","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #6f94b0; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Chapter 1: The sensor_geometry_df Dataframe Analysis</h2>\n\nOnce we got into visualizing the data in the sensor_geometry_df dataframe, let's go over the columns in the dataframe we're in!\n* **sensor_id**: ID specified for the sensors that layed deep into the Antartic ice sheet.\n* **x**, **y**, **z**: The position specified of the sensors.\n\nAnd with that, let's visualize the data specified in the sensor_geometry_df dataframe!","metadata":{}},{"cell_type":"markdown","source":"First, let's visualize the x, y, and z column distributions! We characterize the fig variable to the px module's histogram function for configuring our histogram graph, setting the sensor_geometry_df dataframe as our data for the histogram, the x parameter to the x, y, and z columns for specifying the histogram's x-axes separately, and the marginal parameter to box to plot our box plot in the histogram. With that completed, we exhibit the fig variable figure plot with the show function!","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(sensor_geometry_df, x=\"x\", marginal=\"box\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-11T00:23:38.839092Z","iopub.execute_input":"2023-02-11T00:23:38.840648Z","iopub.status.idle":"2023-02-11T00:23:42.444783Z","shell.execute_reply.started":"2023-02-11T00:23:38.840591Z","shell.execute_reply":"2023-02-11T00:23:42.442444Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.histogram(sensor_geometry_df, x=\"y\", marginal=\"box\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-11T00:23:42.446574Z","iopub.execute_input":"2023-02-11T00:23:42.447428Z","iopub.status.idle":"2023-02-11T00:23:42.643739Z","shell.execute_reply.started":"2023-02-11T00:23:42.447391Z","shell.execute_reply":"2023-02-11T00:23:42.640138Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.histogram(sensor_geometry_df, x=\"z\", marginal=\"box\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-11T00:23:42.645671Z","iopub.execute_input":"2023-02-11T00:23:42.647929Z","iopub.status.idle":"2023-02-11T00:23:42.789021Z","shell.execute_reply.started":"2023-02-11T00:23:42.647890Z","shell.execute_reply":"2023-02-11T00:23:42.786340Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the histogram we generated based on the x column in the sensor_geometry_df dataframe, we noted that the bins in the histogram displayed a nearly symmetrical and bi-modal distribution, in which the ranges 40 to 59.9 and 100 to 119.9 are counted the most, with 240 entities listed, while there are some ranges that have 60 entities counted. And for the box plot plotted on the top of the histogram, the median of the x data column values is 16.99, the first and third quartiles is -224.09 and 224.58, and the interquartile range is 448.67.\n\nMeanwhile on the y column distribution, we envisaged that the bins in this histogram created a mostly symmetrical distribution and somewhat a bi-modal distribution, just like the x column distribution. Nevertheless, the range between -80 to -60.01 in the y data column is counted the most, as there are 300 entities calculated, while there's some ranges in the y column that are counted 60. Thus for the box plot specified on the top of the y column distribution, the median for all data values is -6.055, the first and third quartiles is -283.25 and 228.585, and the interquartile range is 511.610.\n\nLastly for the z column distribution, we found out that this histogram's bins showed almost a uniform distribution, though it was somewhat skewed to the left of the graph. Additionally, the most counted range in the z data column distribution is between -300 and -280.01, as there are 138 data entities listed in here, while the least counted range in the histogram is in between 520 to 539.99, with just 2 entities listed. \n\nSpecifically, the distribution based on data in the x, y, and z columns shows us that they were used for locating the positions of the sensors, as they are collecting the neutrino data deep into the thick layer of ice from Antarctica, since they were found in the abyssal ice. ","metadata":{}},{"cell_type":"markdown","source":"Now let's visualize the scatter plot in x, y, and z perspectives, in Altair! Before that, we have to use the disable_max_rows function from the alt module's data_transformers attribute so that we can plot multiple data without the MaxRowsError. After that, we use the Chart function from the alt module to configure our chart, setting the sensor_geometry_df dataframe as our data for the graph, then configure our scatter plot with the mark_circle function in which we set the size parameter to 60 for configuring the size of the circles in the scatter plot, and then encode them with the encode function for configuring the figure we created, setting the x parameter to the x and y columns separately for the setting the graph's x-axis, the y parameter to the y and z columns for setting the graph's y-axis, the color parameter to the sensor_id column for indicating the color legend for the scatter plot graph, and the tooltip parameter to a list containing the names from the columns x, y (y, z separately), and sensor_id for making the stats appear when hovering the plots, thus on outside from the encode function, we apply the interactive function to make our graph interactive.","metadata":{}},{"cell_type":"code","source":"alt.data_transformers.disable_max_rows()\n\nalt.Chart(sensor_geometry_df).mark_circle(size=60).encode(\n    x=\"x\",\n    y=\"y\",\n    color=\"sensor_id\",\n    tooltip=[\"x\", \"y\", \"sensor_id\"]\n).interactive()","metadata":{"execution":{"iopub.status.busy":"2023-02-11T00:23:42.790523Z","iopub.execute_input":"2023-02-11T00:23:42.790863Z","iopub.status.idle":"2023-02-11T00:23:43.073447Z","shell.execute_reply.started":"2023-02-11T00:23:42.790835Z","shell.execute_reply":"2023-02-11T00:23:43.072449Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"alt.Chart(sensor_geometry_df).mark_circle(size=60).encode(\n    x=\"x\",\n    y=\"z\",\n    color=\"sensor_id\",\n    tooltip=[\"x\", \"z\", \"sensor_id\"]\n).interactive()","metadata":{"execution":{"iopub.status.busy":"2023-02-11T00:23:43.074747Z","iopub.execute_input":"2023-02-11T00:23:43.075047Z","iopub.status.idle":"2023-02-11T00:23:43.663080Z","shell.execute_reply.started":"2023-02-11T00:23:43.075020Z","shell.execute_reply":"2023-02-11T00:23:43.658763Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<!-- <p style=\"text-align: center\"><b>This EDA Simplified Notebook is under beta development, meaning that there's work in progress. It'll be updated shortly while I am working on it! -->","metadata":{}},{"cell_type":"code","source":"alt.Chart(sensor_geometry_df).mark_circle(size=60).encode(\n    x=\"y\",\n    y=\"z\",\n    color=\"sensor_id\",\n    tooltip=[\"y\", \"z\", \"sensor_id\"]\n).interactive()","metadata":{"execution":{"iopub.status.busy":"2023-02-11T00:23:43.664900Z","iopub.execute_input":"2023-02-11T00:23:43.665535Z","iopub.status.idle":"2023-02-11T00:23:44.077269Z","shell.execute_reply.started":"2023-02-11T00:23:43.665387Z","shell.execute_reply":"2023-02-11T00:23:44.074388Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the three scatter plots compiled from three code cells, we visualized that for the x and y scatter plot distribution, it displayed a hexagonal shape as it contained the top outline of the IceCube core detector, in which this scatter plot based on the data in the x and y columns shows the ice top. Meanwhile in the x and z data distribution along with the y and z data distribution, we noticed that we see the deep layer of sensors from the ice sheet, as the IceCube core detector is clearly seen in a cluster of sensors. In other words, these three scatter plots implies us that these are the perspectives of the IceCube observatory.","metadata":{}},{"cell_type":"markdown","source":"And speaking about the three perspectives we visualized in the three scatter plots, let's visualize the whole perspective of the x, y, and z column data with a 3D scatter plot in Plotly! We create our 3d scatter plot by characterizing the fig variable to the px module's scatter_3d function, setting the sensor_geometry_df dataframe as our data input, the x parameter to the x column for configuring the x-axis of the 3d scatter plot, the y-axis to the y column for configuring the y-axis of the 3d scatter plot, the z parameter to the z column for configuring the z-axis of the 3d scatter plot, and the color parameter to the sensor_id variable for arranging the legend indication of the 3d scatter graph. We then update our 3d scatter plot with the update_traces function to the fig figure variable, setting the marker_size parameter to 5 for setting the size of the plot. With all of that completed, we convey our fig variable with the show function.","metadata":{}},{"cell_type":"code","source":"fig = px.scatter_3d(sensor_geometry_df, x='x', y='y', z='z', color='sensor_id')\nfig.update_traces(marker_size=5)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-11T00:23:44.082771Z","iopub.execute_input":"2023-02-11T00:23:44.088697Z","iopub.status.idle":"2023-02-11T00:23:44.489092Z","shell.execute_reply.started":"2023-02-11T00:23:44.088644Z","shell.execute_reply":"2023-02-11T00:23:44.487435Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From this scatter plot graph we created, we visualized the whole perspective of the x, y, and z data columns based on the three scatter plots we plotted separately with Altair. In other words, it displayed the whole outline of the IceCube Neutrino Observatory, containing 60 DOMs in each string of plot, as they are set 125 meters apart.\n\n<center>\n    <img src=\"https://storage.googleapis.com/kaggle-media/competitions/IceCube/icecube_detector.jpg\" width=500>\n    <figcaption style=\"color: gray;\">The whole model of the IceCube Observatory shown above.<figcaption>\n</center>","metadata":{}},{"cell_type":"markdown","source":"And so, we completed our whole analysis in the sensor_geometry_df dataframe! Without further delay, lets visualize the train_df dataframe data with help from Plotly and Altair plotting!","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #6f94b0; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Chapter 2: The train_df Dataframe Analysis</h2>\n\nBefore we visualize the data in the train_df dataframe once we've gotten to the next section from our visualization in the previous section, which is the sensor_geometry_df dataframe, let's explain each column in the train_df dataframe!\n* **batch_id**: ID used for the batch of the event\n* **event_id**: ID used for each neutrino event\n* **(first/last)_pulse_index**: Indexes of the first/last pulse row in the features dataframe belonging to this event\n* **azimuth/zenith**: The angle of the radians of the neutrino in azimuth/zenith.\n\nWith all of the columns in the train_df dataframe explained in detail, let's go onto visualizing the train_df dataframe by each column at a time!","metadata":{}},{"cell_type":"markdown","source":"Before we begin this analysis, we import the two additional modules for plotting, which is the matplotlib module's pyplot submodule as plt and the seaborn modules as sns, since the train_df dataframe contained lots of data, so that we need to preserve RAM and don't commit any SIGILL crashes. ","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns","metadata":{"execution":{"iopub.status.busy":"2023-02-11T00:23:44.490276Z","iopub.execute_input":"2023-02-11T00:23:44.490550Z","iopub.status.idle":"2023-02-11T00:23:45.639706Z","shell.execute_reply.started":"2023-02-11T00:23:44.490525Z","shell.execute_reply":"2023-02-11T00:23:45.636744Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we have the matplotlib and seaborn modules imported, we're ready to visualize the batch_id column from the train_df module! We use the sns module's boxplot function to configure our boxplot chart, setting the data parameter to the train_df dataframe for loading the data for the boxplot chart, and the x parameter to the batch_id column for specifying the x-axes for the boxplot chart.","metadata":{}},{"cell_type":"code","source":"sns.boxplot(data=train_df, x=\"batch_id\")","metadata":{"execution":{"iopub.status.busy":"2023-02-11T00:23:45.642101Z","iopub.execute_input":"2023-02-11T00:23:45.642983Z","iopub.status.idle":"2023-02-11T00:23:55.051059Z","shell.execute_reply.started":"2023-02-11T00:23:45.642952Z","shell.execute_reply":"2023-02-11T00:23:55.047791Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Once we configured our boxplot, we realized that there's loads of batches specified in the data in the train_df dataframe. Nevertheless, the median of the box chart is about 340, the first quartile is about 180 and the third quartile is about 499, and the interquartile range is about 319. In other words, the batch_id is specified for grouping each neutrino event occurances, as it happen frequently in the IceCube Observatory.","metadata":{}},{"cell_type":"markdown","source":"Now let's visualize the first_pulse_index data column in a histogram and the boxplot in two subplots! Before we plot the box graph and the histogram, we characterize fig and ax variables the plt module's subplots function to set up the subplots, setting the 1 and 2 as our number of rows and columns as for the subplot, as well as setting the figsize parameter to 10 by 5 for configuring the width and height of our subplot. Thus, we use the suptitle to the fig variable for setting the title of the subplots, setting it to \"first_pulse_index Distribution\".\n\nNow that we created the subplots, we create our histogram with the sns module's histplot, setting the data parameter to the train_df dataframe's first 10000 rows of data specified by the head function as our data for the histogram, the x parameter to the first_pulse_index column for specifying the histogram's x-axis, the kde parameter to True for showing the kernel density estimation, and the ax parameter to the first slice index of the ax variable (indicated as 0), as well as creating our boxplot with the boxplot function from the sns module, setting the data and x parameter to the same configuration from the histplot configuration, but as for setting the ax parameter, we use the last slice index of the ax variable (indicated as 1). ","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(1, 2, figsize=(10, 5))\nfig.suptitle(\"first_pulse_index Distribution\")\n\nsns.histplot(data=train_df.head(10000), x=\"first_pulse_index\", kde=True, ax=ax[0])\nsns.boxplot(data=train_df.head(10000), x=\"first_pulse_index\", ax=ax[1])","metadata":{"execution":{"iopub.status.busy":"2023-02-11T00:23:55.053329Z","iopub.execute_input":"2023-02-11T00:23:55.053727Z","iopub.status.idle":"2023-02-11T00:23:56.032797Z","shell.execute_reply.started":"2023-02-11T00:23:55.053691Z","shell.execute_reply":"2023-02-11T00:23:56.029440Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For the first_pulse_index distribution, we found out that the bins and kde (kernel-density estimation) in the histogram is mostly uniform, as the highest data counts in this histogram is somewhere in 0.10M and 0.15M, with almost 770 data entities listed, while the least counted data range is around 1.40M, with 80 entities calculated. Meanwhile, in the box plot from the right of the subplot, the median of the data in the first_pulse_index data column is roughly 0.80M, the first and third quartiles is approximately 0.30M and 1.20M, and the interquartile range is around 0.90M. Specifically, the data in the first_pulse_index distriubtion shows us that there are the variant indexes over the pulse rows in the dataframe.","metadata":{}},{"cell_type":"markdown","source":"Now let's visualize the data in the last_pulse_index column with a histogram and the boxplot combined in a subplot! We vaguely follow what we did for visualizing the data in the first_pulse_index data column, but for configuring the fig variable's suptitle function to the graph, we set it to \"last_pulse_index Distribution\" and for the histplot and boxplot functions, we set the x parameter to the train_df dataframe's last_pulse_index column.","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(1, 2, figsize=(10, 5))\nfig.suptitle(\"last_pulse_index Distribution\")\n\nsns.histplot(data=train_df.head(10000), x=\"last_pulse_index\", kde=True, ax=ax[0])\nsns.boxplot(data=train_df.head(10000), x=\"last_pulse_index\", ax=ax[1])","metadata":{"execution":{"iopub.status.busy":"2023-02-11T00:23:56.036215Z","iopub.execute_input":"2023-02-11T00:23:56.039674Z","iopub.status.idle":"2023-02-11T00:23:57.171772Z","shell.execute_reply.started":"2023-02-11T00:23:56.039590Z","shell.execute_reply":"2023-02-11T00:23:57.166312Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Just like the first_pulse_index distribution, we noticed that the kde and distribution of the , as the highest data counts in this histogram is somewhere in 0.10M and 0.15M, with almost 770 data entities listed, while the least counted data range is around 1.40M, with 80 entities calculated. Meanwhile, in the box plot from the right of the subplot, the median of the data in the last_pulse_index data column is roughly 0.80M, the first and third quartiles is approximately 0.30M and 1.20M as it is the same as the visualization in the first_pulse_index dataframe, and the interquartile range is around 0.90M. Specifically, the data in the last_pulse_index distriubtion shows us that both first and last pulse indexes are same and variant over the pulse rows in the dataframe.","metadata":{}},{"cell_type":"markdown","source":"Now let's find the linear regression between the data columns of first_pulse_index and last_pulse_index columns with the seaborn module's lmplot! We use the sns module's lmplot function to create our graph fusion between the scatter plot and the histogram, setting the data parameter to the train_df dataframe's first 10000 rows specified by the head function, the x parameter to the first_pulse_index column for loading the x-axis of the lmplot, the y parameter to the last_pulse_index column for loading the y-axis of the lmplot, and the height parameter to 7 for specifying the height of our lmplot graph.","metadata":{}},{"cell_type":"code","source":"sns.lmplot(data=train_df.head(100), x=\"first_pulse_index\", y=\"last_pulse_index\", height=7)","metadata":{"execution":{"iopub.status.busy":"2023-02-11T00:23:57.176433Z","iopub.execute_input":"2023-02-11T00:23:57.176944Z","iopub.status.idle":"2023-02-11T00:23:58.339665Z","shell.execute_reply.started":"2023-02-11T00:23:57.176915Z","shell.execute_reply":"2023-02-11T00:23:58.332965Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the lmplot based out of the data from the first_pulse_index and the last_pulse_index columns, we visualized how the regression line increases while the scatter plots between the first_pulse_index and the last_pulse_index columns clamped and increased together, although there were some outliers in the regression line, notably four of em. Specifically, the first and last pulse indexes in the dataframe increased whenever the neutrinos pass harmlessly through the ice-layer each second.","metadata":{}},{"cell_type":"markdown","source":"Now let's visualize the azimuth and the zenith data distribution in the subplots with the seaborn's histplot! To get into plotting this out, we characterize the fig and ax variables to the plt module's subplots function, setting the values 1 and 2 as the rows and columns of the subplot, the figsize parameter to 10 by 5 for configuring the subplot's width and height, and the sharey parameter to True for allowing the graphs to share the y-axes together. Furthermore, we configure our title of the subplot with the suptitle function to the fig variable, setting it to \"Azimuth and Zenith Distributions\". \n\nAfter we create the subplots, we configure two histplots with the sns module's histplot function, setting the train_df dataframe's first 10000 rows specified by the head function as our data, the x parameter to the azimuth and zenith data columns separately for the x-axis specification for the histplot, the kde parameter to True for visualizing the kernel-density estimation on the two histplots, and the ax parameter to the ax variable's first and last index separately (specified as 0 and 1). ","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(1, 2, figsize=(10, 5), sharey=True)\nfig.suptitle(\"Azimuth and Zenith Distribution\")\n\nsns.histplot(data=train_df.head(10000), x=\"azimuth\", kde=True, ax=ax[0])\nsns.histplot(data=train_df.head(10000), x=\"zenith\", kde=True, ax=ax[1])","metadata":{"execution":{"iopub.status.busy":"2023-02-11T00:23:58.342452Z","iopub.execute_input":"2023-02-11T00:23:58.343256Z","iopub.status.idle":"2023-02-11T00:23:59.476733Z","shell.execute_reply.started":"2023-02-11T00:23:58.343206Z","shell.execute_reply":"2023-02-11T00:23:59.475604Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see in the Azimuth and Zenith distributions with the kernel-density estimation, we found out that the Azimuth data distribution almost shows a uniform distribution, while the Zenith data distribution shows mostly a symmetrical distribution. Nevertheless, the highest data count in the Azimuth data distribution is 3 with almost 480 data entities counted and the lowest data count is in roughly between 5.1 to 5.2 with about 415 data entities listed, while the highest count in the Zenith distribution is around 1.5 with almost 500 data entities counted and the lowest count is around 3.1 with almost 10 data entities listed. Specifically, the azimuth and zenith data specified in the train_df dataframe hinted us that they were used for gathering the data of the radians of the neutrino in angles.","metadata":{}},{"cell_type":"markdown","source":"Now let's visualize the smooth kernel density estimation between the azimth and zenith data with marginal histplots! We characterize the fig variable to the sns module's JointGrid function to plot out the histograms and kernel-density estimation plots all at once, setting the data parameter to the train_df dataframe for loading the data in the graph, the x parameter to the azimuth column for loading the x-axes for the graph, the y parameter to the zenith column for loading the y-axes for the graph, and the space parameter to 0 for configuring the space value for the graph. We then plot each joint of our fig variable figure graph with the plot_joint function, setting the sns module's kdeplot attribute to plot out the kernel-density estimation plot, as well as the fill parameter to True for filling the area under univariate density curves, the clip parameter to the tuple that contains the first tuple of 0 and 7 and another tuple of 0.0 and 3 for preventing the density evaluation outside of these limits, the thresh parameter to 0 for specifying the lowest iso-proportion level at which to draw the contour line for the graph, the levels parameter to 100 for specifying the number of contour levels to draw contours at, and the cmap parameter to any valid color palatte from Seaborn for specifying the colormap for the graph. Hereafter plotting the joint of our fig variable graph, we plot the marginal graphs with the plot_marginals function to the fig variable graph, setting the sns module's histplot attribute for creating the histogram graph, as well as the color parameter to any hex color for coloring the specified graph, alpha parameter to 1 for adjusting the transparency of the graph, and the bins parameter to 25 for specifying the number of bins in the histogram.","metadata":{}},{"cell_type":"code","source":"fig = sns.JointGrid(data=train_df.head(10000), x=\"azimuth\", y=\"zenith\", space=0)\nfig.plot_joint(sns.kdeplot, fill=True, clip=((0, 7), (0.0, 3.5)), thresh=0, levels=100, cmap=\"cubehelix\")\nfig.plot_marginals(sns.histplot, color=\"#03051A\", alpha=1, bins=25)","metadata":{"execution":{"iopub.status.busy":"2023-02-11T00:23:59.480002Z","iopub.execute_input":"2023-02-11T00:23:59.482402Z","iopub.status.idle":"2023-02-11T00:24:10.086629Z","shell.execute_reply.started":"2023-02-11T00:23:59.482332Z","shell.execute_reply":"2023-02-11T00:24:10.082860Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From what we saw based on the relations between the azimuth and the zenith data, we found out that the histplots on the x and y axes of the JointGrid is similar to the ones we plotted separately, but on the kernel-density estimation plot in the JointGrid graph, we envisaged how it almost fills up the kde graph, as we noted how we saw the bright areas in the middle of the graph. Specifically, the zenith and azimuth data is more related to each other when it was each used to collect the data of the neutrino's radians in angles.","metadata":{}},{"cell_type":"markdown","source":"And for all of that, we completed most of our analysis of the train_df dataframe! And for looking forward in our analysis journey, let's visualize the data in one of the batches of parquet files specified in the train folder.","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #6f94b0; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Chapter 3: The Batch Analysis</h2>\n\nAs we got into visualizing the batches specified in the competition data, we found out that they are parquet files too just like the train_df dataframe. Without a doubt, we characterize the batch_df dataframe to use the read_parquet function from the pd module to read out any specified batch paraquet in the train folder. With that completed, we use the head function to display the first five rows in the batch_df dataframe.","metadata":{}},{"cell_type":"code","source":"batch_df = pd.read_parquet(\"/kaggle/input/icecube-neutrinos-in-deep-ice/train/batch_124.parquet\")\nbatch_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-11T00:24:10.088603Z","iopub.execute_input":"2023-02-11T00:24:10.089228Z","iopub.status.idle":"2023-02-11T00:24:14.743276Z","shell.execute_reply.started":"2023-02-11T00:24:10.089181Z","shell.execute_reply":"2023-02-11T00:24:14.737723Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Once we finished reading out any specific parquet batch into our dataframe we created as well as displaying it, we found out that there are 5 columns of data in each batch in several parquet files in the train folder. With that, let's list out each column in a specific batch parquet!\n* **event_id**: ID specified for the specific event.\n* **sensor_id**: ID specified for the sensors underneath the ice sheet.\n* **time**: How long the neutrinos were detected by the sensors.\n* **charge**: The estimate amount of the charge in the neutrinos.\n* **auxiliary**: Determines if the pulse is not fully digitalized or was contributed to the trigger decision.","metadata":{}},{"cell_type":"markdown","source":"Now that we completed explaining the data columns, let's visualize the sensor_id column, into the histogram by Plotly! We define the fig variable to the px module's histogram function, setting the batch_df dataframe as our data input for the histogram, the x parameter to the sensor_id column for specifying the x-axes of the histogram, and the marginal parameter to box for placing the box-plot on the top of the histogram. With that completed, we show our histogram graph with the show function!","metadata":{}},{"cell_type":"code","source":"batch_df[\"sensor_id\"] = batch_df[\"sensor_id\"].astype(int)\n\nfig = px.histogram(batch_df.head(10000), x=\"sensor_id\", marginal=\"box\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-11T00:24:14.745342Z","iopub.execute_input":"2023-02-11T00:24:14.745769Z","iopub.status.idle":"2023-02-11T00:24:15.782710Z","shell.execute_reply.started":"2023-02-11T00:24:14.745734Z","shell.execute_reply":"2023-02-11T00:24:15.777218Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From this histogram as well as the box chart above, we visualized that the bins in the histogram graph shows a mostly uniform distribution, as there are small and big peaks of data in the histogram. Nevertheless, the highest data count in this sensor_id distribution is under 100 to 199, with 1665 data entities counted, while the range between 2300 and 2399 is counted the least, as there are 46 entities calculated. In other words, the data distribution of the sensor_id data columns explains how the id specified in the sensors under the sheet of ice detects every neutrino passing to them under each specific event batch.","metadata":{}},{"cell_type":"markdown","source":"Now let's visualize the time column distribution in a histogram and the box plot with help from Altair! We characterize the one variable graph to use the alt module's Chart function to configure our chart in Altair, setting the batch_df dataframe's first 10000 rows with the head function as our data for the graph, then place the mark_bar function to create the bars in the histogram graph as well as placing our graph's configuration with the encode function, setting the alt module's X function for placing the x-axes in the graph, containing the time column for specifying the column of the data, as well as the bin parameter to True for displaying the histogram along with the y parameter to the count function encased in strings for counting the values of the specified data in the histogram, as well as characterizing the two variable to do the same thing for arranging the Chart function from the alt module, but this time, we apply the mark_boxplot function for creating our boxplot, setting the extent parameter to \"min-max\" for configuring the minimum and maximum range for the boxplot thus configuring our boxplot graph with the encode function, setting the x parameter to the time column for configuring the x-axes of the boxplot, along with editing the graph's properties with the properties function, setting the height to 300 for configuring the height of the boxplot. Finally, we merge our two graphs into our subplot with the concat function from the alt module, putting the one and two variables inside of it.","metadata":{}},{"cell_type":"code","source":"one = alt.Chart(batch_df.head(10000)).mark_bar().encode(\n    alt.X(\"time\", bin=True),\n    y=\"count()\"\n)\n\ntwo = alt.Chart(batch_df.head(10000)).mark_boxplot(extent='min-max').encode(\n    x='time',\n).properties(height=300)\n\nalt.concat(one, two)","metadata":{"execution":{"iopub.status.busy":"2023-02-11T00:24:15.792302Z","iopub.execute_input":"2023-02-11T00:24:15.793666Z","iopub.status.idle":"2023-02-11T00:24:16.440983Z","shell.execute_reply.started":"2023-02-11T00:24:15.793148Z","shell.execute_reply":"2023-02-11T00:24:16.436441Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From our Altair-made subplot graphs consisting over the distribution of the time column in the histogram and the boxplot, we noticed that in the histogram chart on the left of the subplots, we espied that the bins of the chart displayed a left skew in the chart, but specifically, the highest counts in this histogram is in between 10K and 15K, with approximately 6995 data entities listed, while the least counted is in the range from 25K to 30K, as there are about 200 data entities counted. \n\nMeanwhile in the box plot, we visualized that the line of distance from the interquartile box range to the maximum of the time data column is way longer than the line of distance from the interquartile box range to the minimum of the time data column. Not only that, we found out that the first quartile of the box plot is 10214.75, the median is 11637.5, the third quartile is 13380.25, and the interquartile range is 3165.5. Specifically for the two graphs shown above, the data in the time column hinted us that there are variations of time in every neutrino passing towards the IceCube Neutrino Observatory.","metadata":{}},{"cell_type":"markdown","source":"Now let's visualize the charge data column with our histogram as well as the box plot with Plotly! We characterize the fig variable to the px module's histogram function to configure our histogram diagram, setting the batch_df dataframe's first 10000 rows specified by the head function as our data for the histogram, as well as the x parameter to the charge data column for specifying the x-axes of our histogram, and the marginal parameter to box for placing our box plot on the top of our histogram. Finally, we apply the show function to the fig variable figure we configured for exhibiting our graph.","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(batch_df.head(10000), x=\"charge\", marginal=\"box\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-11T00:24:16.447786Z","iopub.execute_input":"2023-02-11T00:24:16.448287Z","iopub.status.idle":"2023-02-11T00:24:16.727309Z","shell.execute_reply.started":"2023-02-11T00:24:16.448251Z","shell.execute_reply":"2023-02-11T00:24:16.719993Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From our data distribution based on the charge data column, we found out that our histogram displayed a skew towards the left of our diagram. In other words, we visualized that the highest data count in this charge histogram is between 0.8 and 0.99, with 1837 data entities listed, while we magnified the single counted bins of data range on the right of our histogram. And as for the boxplot we compiled on the top of our histogram diagram, we spotted a lot of outliers throughout our box graph, but the box part in the box plot is too narrow and small that we must zoom in to see the box part clearly. And speaking about the box part, the median of the charge data column is 1.025, the first quartile is 0.725, the third quartile is 1.375, and the interquartile range is 0.65. Furthermore, the data in the charge column shows that the pulse of light from the neutrinos mostly measure between from zero to two once they passed towards the IceCube Observatory.","metadata":{}},{"cell_type":"markdown","source":"Once we visualized the data distribution of the time and charge columns, let's visualize the data relationships of them by the sensor_id column in a scatter plot! We characterize our chart with the alt module's Chart function, setting the batch_df dataframe's first 10000 rows specified by the head function as our data for the chart, as well as creating our scatter plot with the mark_circle function in which we configured the size parameter to 60 for setting the size of a circle, thus configuring the heatmap diagram with the encode function, setting the x parameter to the time column for our scatter plot's x-axes, the y parameter to the charge column for our scatter plot's y-axes, the color parameter to the sensor_id column for specifying the color range for the scatter plot, and the tooltip parameter to a list containing the columns of time, charge, and sensor_id for making a plot description when hovered.","metadata":{}},{"cell_type":"code","source":"alt.Chart(batch_df.head(10000)).mark_circle(size=60).encode(\n    x=\"time\",\n    y=\"charge\",\n    color=\"sensor_id\",\n    tooltip=[\"time\", \"charge\", \"sensor_id\"]\n)","metadata":{"execution":{"iopub.status.busy":"2023-02-11T00:24:16.729565Z","iopub.execute_input":"2023-02-11T00:24:16.730248Z","iopub.status.idle":"2023-02-11T00:24:17.296183Z","shell.execute_reply.started":"2023-02-11T00:24:16.730210Z","shell.execute_reply":"2023-02-11T00:24:17.288037Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the scatter plot we created, we found out that most plots based on the data in the charge and time columns were plotted horizontally at the bottom of the scatter plot than plotted vertically at the left of the scatter plot. Specifically, the number of data plotted at the bottom in the scatter plot shows us that most the neutrino's light pulses that are ranged from 0 to 2 were related towards the variant time specification as the neutrinos passed towards the IceCube Observatory.","metadata":{}},{"cell_type":"markdown","source":"Once we distributed and visualized the time and charge data, let's visualize the data in the auxiliary column with a pie chart in Plotly! Before we began plotting this diagram, we inherit our dataframe, auxiliary to the first 10000 rows of the batch_df dataframe's auxiliary data specified by the head function and count their values with the value_counts function. After that, we configure our Plotly graph by defining our fig variable to the pie function from the px module, setting the auxiliary dataframe as our data for the pie chart, the names parameter to the auxiliary dataframe's indexes specified by the index attribute for specifying the names of the pie chart, and the values to the values of the auxiliary dataframe specified by the values attribute for configuring the values of the pie chart. Finally, we apply the show function to the fig variable figure for displaying our pie chart.","metadata":{}},{"cell_type":"code","source":"auxiliary = batch_df.auxiliary.head(10000).value_counts()\n\nfig = px.pie(auxiliary, names=auxiliary.index, values=auxiliary.values)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-11T00:24:25.354682Z","iopub.execute_input":"2023-02-11T00:24:25.355621Z","iopub.status.idle":"2023-02-11T00:24:25.482596Z","shell.execute_reply.started":"2023-02-11T00:24:25.355560Z","shell.execute_reply":"2023-02-11T00:24:25.481531Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see for the pie chart, we found out that there are more data in the auxiliary column that are marked as false than being marked as true. In other words, the total counts of data that is labeled as false is 6230, while the other data that is labeled as true is 3770. Specifically, the number of data listed as false in the batch_df dataframe's auxiliary column hinted us that most pulses contributed to the trigger decision as well as being fully digitalized in most time occurances from the specified event.","metadata":{}},{"cell_type":"markdown","source":"And with all of the columns visualized with our created graphs, we completed our data analysis of the batch_df dataframe that is generated from the batch parquet files in the train folder!","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #6f94b0; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Conclusion</h2>\n\nFrom our data analysis journey based on the data we visualized in the sensor_geometry_df, train_df, and batch_df dataframes, we found out that the data gathered in the IceCube Observatory's sensors based on the passing neutrinos has a range on zero to two on light pulses as they absolutely travel at any speed at all in any time occurences in each event as well as mapping out the trail of a straight neutrino based from the pulse indexes we visualized with regression. And now that we visualized the data in the IceCube Neutrino Competition, we will expect to see scientists extend their understanding over exploding stars, bursting gamma-rays, and even unravel the unsolved mysteries of catalysmic phenomena like black holes, neutron stars, and the notable neutrinos. Looking forward from that competition we're in, the hunt for neutrinos rests on our study of the frontier science physics in around the observatories from IceCube to Kamiokande and COBRA!","metadata":{}}]}