{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# IceCube❄️: Inference🚀  with RandomForest regression\n\nThe goal of this competition is to predict a neutrino particle’s direction. You will develop a model based on data from the \"IceCube\" detector, which observes the cosmos from deep within the South Pole ice.\n\n**Notattion:** in my notebook i will use Helped Notebook from other contrubitors to help focus on Design the model effectively \n\n#### **Resources** \nSee EDA: [**IceCube🧊: Neutrino🎆EDA & 3D🔭interactive viewer**](https://www.kaggle.com/code/jirkaborovec/icecube-neutrino-eda-3d-interactive-viewer)\n\nsee get hands-on Data:[**日本語/Eng🧊IceCube EDA: Understanding train data**](https://www.kaggle.com/code/utm529fg/eng-icecube-eda-understanding-train-data)\n\ninspiration notebook: [**IceCube🧊: Neutrino🎆 baseline XGBoost🚀Regression**](https://www.kaggle.com/code/jirkaborovec/icecube-neutrino-baseline-xgboost-regression)\n\n**Tutorials and Tools used:**\n\ncheck **RandomizedSearchCV-**  [**RandomizedSearchCV**](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.RandomizedSearchCV.html) \n\n**Hyper-Tuing the parameters**\ncheck - [**RandomForestClassifier**](https://scikit-learn.org/stable/modules/generated/sklearn.ensemble.RandomForestClassifier.html) \n\n**the model Regression**\ncheck - [**decomposition.PCA**](https://scikit-learn.org/stable/modules/generated/sklearn.decomposition.PCA.html) **Dimessionality Reduction**\n","metadata":{}},{"cell_type":"markdown","source":"#### The Task Cover in this notebook \n- Following all the shared notebook in competition by awosome contrubitors out there i will split the problem of this competition in sub-Tasks :\n   * **Read the data and understand it**\n      1. Data Description \n      2. define the outlines of each Colums in data \n   * **Virtualization**\n      1. Convert (azimuth,zenith) into Sphrice Corredinats \n      2. plot Senors_ID at space plan \n      3. show the Heatmap of azimuth,zenith \n   * Extracting Features (**Feature engineering**)\n   * implment Randorm Forset and Optimization step using **GridSearchCV** Hyper-tune Parameters ","metadata":{}},{"cell_type":"code","source":"%matplotlib inline\nimport os\nimport glob\nimport math\nimport numpy as np\nimport pandas as pd\nfrom tqdm.auto import tqdm\nimport matplotlib.pyplot as plt\nimport plotly.express as px\nimport plotly.graph_objects as go\nfrom plotly.subplots import make_subplots\nfrom tqdm.notebook import tqdm\n\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.decomposition import PCA\nfrom sklearn.pipeline import Pipeline\nfrom sklearn.ensemble import RandomForestRegressor \nfrom sklearn.multioutput import MultiOutputRegressor\nfrom sklearn.model_selection import RandomizedSearchCV\nfrom sklearn.preprocessing import MinMaxScaler\nfrom sklearn import metrics\nfrom sklearn.metrics import r2_score, make_scorer\nfrom sklearn.model_selection import RepeatedKFold\nfrom sklearn.metrics import mean_squared_error, mean_absolute_error\n\nPATH_DATASET = \"/kaggle/input/icecube-neutrinos-in-deep-ice\"","metadata":{"execution":{"iopub.status.busy":"2023-02-27T18:13:32.009518Z","iopub.execute_input":"2023-02-27T18:13:32.010037Z","iopub.status.idle":"2023-02-27T18:13:35.942622Z","shell.execute_reply.started":"2023-02-27T18:13:32.009990Z","shell.execute_reply":"2023-02-27T18:13:35.941246Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Read the data and understand it**","metadata":{}},{"cell_type":"code","source":"meta_data_test = pd.read_parquet(os.path.join(PATH_DATASET, \"test_meta.parquet\"))\nprint(f\"length: {len(meta_data_test)}\")\nmeta_data_test.head()\ndel meta_data_test","metadata":{"execution":{"iopub.status.busy":"2023-02-27T18:13:35.944894Z","iopub.execute_input":"2023-02-27T18:13:35.945310Z","iopub.status.idle":"2023-02-27T18:13:36.103299Z","shell.execute_reply.started":"2023-02-27T18:13:35.945271Z","shell.execute_reply":"2023-02-27T18:13:36.101920Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Understanding of Data Columns\n\n**[train/test]_meta.parquet**\n\n- **batch_id** (int): the ID of the batch the event was placed into.\n- **event_id** (int): the event ID.\n- **[first/last]_pulse_index** (int): index of the first/last row in the features dataframe belonging to this event.\n- **[azimuth/zenith]** (float32): the [azimuth/zenith] angle in radians of the neutrino. A value between 0 and 2*pi for the azimuth and 0 and pi for zenith. The target columns. Not provided for the test set. The direction vector represented by zenith and azimuth points to where the neutrino came from.","metadata":{}},{"cell_type":"code","source":"meta_data_train = pd.read_parquet(os.path.join(PATH_DATASET, \"train_meta.parquet\"))\nprint(f\"total events: {len(meta_data_train)}\")\nprint(f\"nb batches: {len(meta_data_train['batch_id'].unique())}\")\nmeta_data_train.head(5)","metadata":{"execution":{"iopub.status.busy":"2023-02-27T18:13:36.104952Z","iopub.execute_input":"2023-02-27T18:13:36.105301Z","iopub.status.idle":"2023-02-27T18:14:25.285183Z","shell.execute_reply.started":"2023-02-27T18:13:36.105266Z","shell.execute_reply":"2023-02-27T18:14:25.283790Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### outline in depth training/test data\n\n**[train/test]/batch_[n].parquet** Each batch contains tens of thousands of events. Each event may contain thousands of pulses, each of which is the digitized output from a photomultiplier tube and occupies one row.\n\n- **event_id** (int): the event ID. Saved as the index column in parquet.\n- **time** (int): the time of the pulse in nanoseconds in the current event time window. The absolute time of a pulse has no relevance, and only the relative time with respect to other pulses within an event is of relevance.\n- **sensor_id** (int): the ID of which of the 5160 IceCube photomultiplier sensors recorded this pulse.\n- **charge** (float32): An estimate of the amount of light in the pulse, in units of photoelectrons (p.e.). A physical photon does not exactly result in a measurement of 1 p.e. but rather can take values spread around 1 p.e. As an example, a pulse with charge 2.7 p.e. could quite likely be the result of two or three photons hitting the photomultiplier tube around the same time. This data has float16 precision but is stored as float32 due to limitations of the version of pyarrow the data was prepared with.\n- **auxiliary** (bool): If True, the pulse was not fully digitized, is of lower quality, and was more likely to originate from noise. If False, then this pulse was contributed to the trigger decision and the pulse was fully digitized.","metadata":{}},{"cell_type":"code","source":"df_test = pd.read_parquet(os.path.join(PATH_DATASET, \"test/batch_661.parquet\"))\nprint(f\"length: {len(df_test)}\")\nprint(f\"events: {len(df_test.index.unique())}\")\ndisplay(df_test.head())\ndel df_test","metadata":{"execution":{"iopub.status.busy":"2023-02-27T18:14:25.288872Z","iopub.execute_input":"2023-02-27T18:14:25.290036Z","iopub.status.idle":"2023-02-27T18:14:25.329194Z","shell.execute_reply.started":"2023-02-27T18:14:25.289977Z","shell.execute_reply":"2023-02-27T18:14:25.327659Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* **sensor_geometry.csv description**\n\n    * The **x, y**, and **z positions** for each of the **5160 IceCube sensors**.\n    * The row index corresponds to the sensor_idx feature of pulses.\n    * The **x, y**, and **z coordinates** are in units of meters, **with the origin at the center of the IceCube detector**.\n    * The coordinate system is right-handed, and the z-axis points upwards when standing at the South Pole.</br>\n   **Notation** \n       - You can **convert from these coordinates to azimuth and zenith with the following formulas (here the vector (x,y,z) is normalized)**\n","metadata":{}},{"cell_type":"code","source":"meta_data_sensor_position = pd.read_csv(PATH_DATASET + \"/sensor_geometry.csv\")\nprint(f\"lengh of data: {len(meta_data_sensor_position.sensor_id)}\")\nmeta_data_sensor_position.head(5)\n","metadata":{"execution":{"iopub.status.busy":"2023-02-27T18:14:25.330636Z","iopub.execute_input":"2023-02-27T18:14:25.330981Z","iopub.status.idle":"2023-02-27T18:14:25.368950Z","shell.execute_reply.started":"2023-02-27T18:14:25.330947Z","shell.execute_reply":"2023-02-27T18:14:25.367481Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Virtualization**\n\n\n**azimuth and zenith**\n   * The solar azimuth and solar zenith express the position of the sun. The solar azimuth is the angle of the direction of the sun measured clockwise north from the horizon. The solar zenith is the angle measured from the local zenith and the line of sight of the sun.","metadata":{}},{"cell_type":"code","source":"plt.hist2d(meta_data_train[\"azimuth\"], meta_data_train[\"zenith\"], bins=(50, 50), cmap=plt.cm.jet)\nplt.xlabel('azimuth'), plt.ylabel('zenith')\nplt.colorbar()","metadata":{"execution":{"iopub.status.busy":"2023-02-27T18:14:25.371135Z","iopub.execute_input":"2023-02-27T18:14:25.372078Z","iopub.status.idle":"2023-02-27T18:14:45.068806Z","shell.execute_reply.started":"2023-02-27T18:14:25.372016Z","shell.execute_reply":"2023-02-27T18:14:45.066919Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"meta_data_train_ = meta_data_train[meta_data_train['batch_id'] == 10]\nprint(f\"batch events: {len(meta_data_train_)}\")\nplt.hist2d(meta_data_train_[\"azimuth\"], meta_data_train_[\"zenith\"], bins=(50, 50), cmap=plt.cm.jet)\nplt.xlabel('azimuth'), plt.ylabel('zenith')\nplt.colorbar()","metadata":{"execution":{"iopub.status.busy":"2023-02-27T18:14:45.071476Z","iopub.execute_input":"2023-02-27T18:14:45.072476Z","iopub.status.idle":"2023-02-27T18:14:46.210513Z","shell.execute_reply.started":"2023-02-27T18:14:45.072430Z","shell.execute_reply":"2023-02-27T18:14:46.209048Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# convert to x, y, z coordinate\nmeta_data_train[\"x\"] = np.cos(meta_data_train[\"azimuth\"]) * np.sin(meta_data_train[\"zenith\"])\nmeta_data_train[\"y\"] = np.sin(meta_data_train[\"azimuth\"]) * np.sin(meta_data_train[\"zenith\"])\nmeta_data_train[\"z\"] = np.cos(meta_data_train[\"zenith\"])\n\n# sampling 10000\ndf_train_sample = meta_data_train.sample(10000)\n\n# 3d plot\nfig = px.scatter_3d(df_train_sample, x='x', y='y', z='z', color='z', opacity=1)\nfig.update_traces(marker_size=1)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-27T18:14:46.212146Z","iopub.execute_input":"2023-02-27T18:14:46.212684Z","iopub.status.idle":"2023-02-27T18:15:25.024162Z","shell.execute_reply.started":"2023-02-27T18:14:46.212479Z","shell.execute_reply":"2023-02-27T18:15:25.022824Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n***Delay between first and last pulse¶***\n","metadata":{}},{"cell_type":"code","source":"meta_data_train[\"delay_pulse_index\"] = meta_data_train[\"last_pulse_index\"] - meta_data_train[\"first_pulse_index\"]\n_ = plt.hist(meta_data_train[\"delay_pulse_index\"], bins=100, log=True)\nplt.ylabel('count cases'), plt.xlabel('duration')\nplt.grid()\n","metadata":{"execution":{"iopub.status.busy":"2023-02-27T18:15:25.025927Z","iopub.execute_input":"2023-02-27T18:15:25.026291Z","iopub.status.idle":"2023-02-27T18:15:29.644722Z","shell.execute_reply.started":"2023-02-27T18:15:25.026254Z","shell.execute_reply":"2023-02-27T18:15:29.643507Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n**Browse geometry**\n\nThe x, y, and z positions for each of the 5160 IceCube sensors. \nThe row index corresponds to the sensor_idx feature of pulses. The x, y, and z coordinates are in units of meters, with the origin at the center of the IceCube detector. The coordinate system is right-handed, and the z-axis points upwards when standing at the South Pole.\n","metadata":{}},{"cell_type":"code","source":"fig = px.scatter_3d(meta_data_sensor_position, x='x', y='y', z='z', color='sensor_id', opacity=0.7)\nfig.update_traces(marker_size=2)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-27T18:15:29.649319Z","iopub.execute_input":"2023-02-27T18:15:29.649716Z","iopub.status.idle":"2023-02-27T18:15:29.732892Z","shell.execute_reply.started":"2023-02-27T18:15:29.649679Z","shell.execute_reply":"2023-02-27T18:15:29.731800Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"meta_data_train[meta_data_train['event_id'].isin([46528394, 2135637939, 2084362251])]","metadata":{"execution":{"iopub.status.busy":"2023-02-27T18:15:29.734344Z","iopub.execute_input":"2023-02-27T18:15:29.734996Z","iopub.status.idle":"2023-02-27T18:15:36.149538Z","shell.execute_reply.started":"2023-02-27T18:15:29.734947Z","shell.execute_reply":"2023-02-27T18:15:36.148266Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Simulate the IceCube Eniveremenet\n**Notation:** we will simulate per one Batch because each batch contain ten of thouns **Event_ID**</br>\n* in this Simulation show how does Telescop Detect the Neutons Direction by Projection</br>\n\n    **Mechanism Description:**</br>\n    IceCube is a neutrino observatory located at the South Pole. It detects neutrinos by observing the light produced when neutrinos interact with the ice surrounding the observatory. However, networks of telescopes worldwide can also play a role in the detection of neutrinos.\n\n    When a high-energy neutrino enters the Earth's atmosphere, it can produce a cascade of particles, including muons, that travel at nearly the speed of light. These particles can emit a faint blue light, known as Cherenkov radiation, as they move through the atmosphere. This light can be detected by large arrays of optical telescopes, such as the High Altitude Water Cherenkov (HAWC) observatory in Mexico or the Imaging Atmospheric Cherenkov Telescope (IACT) arrays in various locations around the world.\n\n    If a neutrino enters the Earth and interacts with the ice in the IceCube detector, it can produce a high-energy muon that will travel through the ice, emitting Cherenkov radiation as it goes. This light can be detected by the photomultiplier tubes (PMTs) in the IceCube detector.\n\n    By combining the information from the telescopes that detect the Cherenkov radiation in the atmosphere with the information from the IceCube detector, scientists can determine the direction and energy of the neutrino. This information can be used to study the sources of the neutrinos and to learn more about the high-energy astrophysical processes that produce them.\n\n    In summary, networks of telescopes worldwide can detect the Cherenkov radiation produced by high-energy particles, including muons produced by neutrinos interacting with the Earth's atmosphere. When combined with data from the IceCube detector, this information can be used to study the properties and sources of neutrinos.","metadata":{}},{"cell_type":"code","source":"train = pd.read_parquet(os.path.join(PATH_DATASET, \"train/batch_101.parquet\"))\nprint(f\"length: {len(train)}\")\nprint(f\"events: {len(train.index.unique())}\")\ntrain.head()\n","metadata":{"execution":{"iopub.status.busy":"2023-02-27T18:15:36.151232Z","iopub.execute_input":"2023-02-27T18:15:36.151726Z","iopub.status.idle":"2023-02-27T18:15:41.583101Z","shell.execute_reply.started":"2023-02-27T18:15:36.151676Z","shell.execute_reply":"2023-02-27T18:15:41.581617Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def charge_center(event):\n    evt_sub = event[~event['auxiliary']][['x', 'y', 'z', 'charge']]\n    evt_sub.loc[:, 'coef'] = evt_sub['charge'] / evt_sub['charge'].sum()\n    # display(evt_sub)\n    for c in ['x', 'y', 'z']:\n        evt_sub.loc[:, c] *= evt_sub['coef']\n    cx, cy, cz = evt_sub[['x', 'y', 'z']].sum().values\n    return cx, cy, cz\n\ndef direction(meta):\n    azimuth = meta[\"azimuth\"]\n    zenith = meta[\"zenith\"]\n    dx = math.sin(zenith) * math.cos(azimuth)\n    dy = math.sin(zenith) * math.sin(azimuth)\n    dz = math.cos(zenith)\n    return dx, dy, dz","metadata":{"execution":{"iopub.status.busy":"2023-02-27T18:15:41.584687Z","iopub.execute_input":"2023-02-27T18:15:41.585018Z","iopub.status.idle":"2023-02-27T18:15:41.594857Z","shell.execute_reply.started":"2023-02-27T18:15:41.584984Z","shell.execute_reply":"2023-02-27T18:15:41.593413Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def draw_subplot(fig, i, evt, sensors, cx, cy, cz, dx, dy, dz, scale=400):\n    # sensors as background\n    fig.add_trace(\n        go.Scatter3d(\n            x=sensors['x'], y=sensors['y'], z=sensors['z'], \n            mode='markers', marker=dict(size=1, color=\"black\"), opacity=0.2\n        ), row=(i+1), col=1)\n    # sensors reading\n    fig.add_trace(\n        go.Scatter3d(\n            x=evt['x'], y=evt['y'], z=evt['z'], opacity=0.8,\n            mode='markers', marker=dict(\n                size=evt['charge'] * 15,\n                color=evt['time'],\n                colorscale='sunsetdark',\n            )\n        ), row=(i+1), col=1)\n    # direction from metad data\n    fig.add_trace(\n        go.Scatter3d(\n            x=[cx - dx * scale, cx + dx * scale],\n            y=[cy - dy * scale, cy + dy * scale],\n            z=[cz - dz * scale, cz + dz * scale],\n            opacity=0.8, mode='lines', line=dict(color='red', width=3)\n        ), row=(i+1), col=1)","metadata":{"execution":{"iopub.status.busy":"2023-02-27T18:15:41.596669Z","iopub.execute_input":"2023-02-27T18:15:41.597933Z","iopub.status.idle":"2023-02-27T18:15:41.615106Z","shell.execute_reply.started":"2023-02-27T18:15:41.597888Z","shell.execute_reply":"2023-02-27T18:15:41.612792Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def show_event(event_id=46528394, data=train, sensors=meta_data_sensor_position, metadata=meta_data_train):\n    meta = dict(metadata[metadata[\"event_id\"]==event_id].iloc[0])\n    display(meta)\n    event = data[data.index == event_id]\n    event = event.merge(sensors, on=\"sensor_id\")\n    event.loc[:, 'charge'] /= event['charge'].max()\n    dx, dy, dz = direction(meta)\n    cx, cy, cz = charge_center(event)\n    #print(cx, cy, cz)\n    \n    auxiliaries = [False, True]\n    fig = make_subplots(\n        rows=2, specs=[[{'type': 'scene'}], [{'type': 'scene'}]],\n        subplot_titles=[f\"auxiliary={aux}\" for aux in auxiliaries],\n        vertical_spacing=0.05,\n    )\n    for i, aux in enumerate(auxiliaries):\n        evt_ = event[event['auxiliary'] == aux]\n        draw_subplot(fig, i, evt_, sensors, cx, cy, cz, dx, dy, dz)\n    fig.update_layout(\n        height=800, width=600, showlegend=False,\n        title_text=f\"Event #{event_id}\"\n        f\" / azimuth={meta['azimuth']:0.3}; zenith={meta['zenith']:0.3}\\n\"\n        f\"-> x={dx:0.2}; y={dy:0.2}; z={dz:0.2}\",\n    )\n    # fig = px.scatter_3d(event, x='x', y='y', z='z', size=\"charge\", color=\"auxiliary\",\n    #     opacity=0.8, title=f\"Event #{event_id}\\n -> x={x_}; y={y_}; z={z_}\")\n    return fig\n\nshow_event().show()","metadata":{"execution":{"iopub.status.busy":"2023-02-27T18:15:41.616971Z","iopub.execute_input":"2023-02-27T18:15:41.617328Z","iopub.status.idle":"2023-02-27T18:15:42.138995Z","shell.execute_reply.started":"2023-02-27T18:15:41.617292Z","shell.execute_reply":"2023-02-27T18:15:42.136811Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Extracting Features **Feature engineering**\n* Steps \n     1. **Transfomation training data**\n        - converting data where single row is one event so sensors are columns\n        - **NOTE:** avoid usung pandas as it will blow memory\n        - **NOTE:** use just single batch, maybe depending on memory consimptions we can use more...\n        - **Notation:** since that the data is contain a lot of features to work on them , in **our Case we will take Input/Target as follow** \n           - **Input** (Event_ID , Number_Sensors)\n           - **Target** (Event_ID . Angels)\n        \n  2. **Data preprocessing:**  \n      to reduce amount of data Point and extract only most informativeffeatures which our **dataTraning** has , we will implement to approaches are :</br>\n       1. **PCA** \n           (Principal Component Analysis) : is a technique used in data analysis and machine learning for dimensionality reduction. It works by identifying the most important features or variables in a dataset and transforming the data into a lower-dimensional space that retains most of the important information. \n       2. **MinMaxScaler:**\n          is a data preprocessing technique used in machine learning to scale the features of a dataset to a specified range, usually between 0 and 1. It works by subtracting the minimum value of each feature and then dividing by the range (i.e., the difference between the maximum and minimum values).     \n  ","metadata":{}},{"cell_type":"code","source":"def transform_batch(path_batch_parquet, nb_sensors=5160):\n    df = pd.read_parquet(path_batch_parquet)[:1000000]\n    # df[df['auxiliary']]['charge'] /= 10.\n    data = np.zeros((len(df.index.unique()), nb_sensors), dtype=np.float16)\n    event_ids = []\n    for i, (idx, dfg) in tqdm(enumerate(df.groupby(level=0))):\n        event_ids.append(idx)\n        data[i, dfg['sensor_id']] = dfg['charge']\n    # df = pd.DataFrame(data, index=event_ids, columns=[f\"sid-{i}\" for i in range(nb_sensors)])\n    return event_ids, data","metadata":{"execution":{"iopub.status.busy":"2023-02-27T18:15:42.141110Z","iopub.execute_input":"2023-02-27T18:15:42.142121Z","iopub.status.idle":"2023-02-27T18:15:42.155677Z","shell.execute_reply.started":"2023-02-27T18:15:42.142033Z","shell.execute_reply":"2023-02-27T18:15:42.153596Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"event_ids, data = transform_batch(os.path.join(PATH_DATASET, \"train/batch_10.parquet\"))\nprint(f\"data size: {data.shape}\")\n\nmeta_train_ = meta_data_train[meta_data_train['event_id'].isin(event_ids)]\nmeta_train_ = dict(zip(meta_train_['event_id'].values, meta_train_[[\"azimuth\", \"zenith\"]].values.tolist()))\nprint(f\"LUT size: {len(meta_train_)}\")\nangles = np.array([meta_train_[eid] for eid in event_ids], dtype=np.float16)\nprint(f\"angles size: {angles.shape}\")","metadata":{"execution":{"iopub.status.busy":"2023-02-27T18:15:42.157658Z","iopub.execute_input":"2023-02-27T18:15:42.158218Z","iopub.status.idle":"2023-02-27T18:15:49.250927Z","shell.execute_reply.started":"2023-02-27T18:15:42.158165Z","shell.execute_reply":"2023-02-27T18:15:49.249687Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train, X_test, y_train, y_test = train_test_split(data, angles, train_size=0.8)\ndel data, angles","metadata":{"execution":{"iopub.status.busy":"2023-02-27T18:15:49.252725Z","iopub.execute_input":"2023-02-27T18:15:49.253133Z","iopub.status.idle":"2023-02-27T18:15:49.299676Z","shell.execute_reply.started":"2023-02-27T18:15:49.253097Z","shell.execute_reply":"2023-02-27T18:15:49.298521Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"preprocess = Pipeline([\n    ('scaler',MinMaxScaler()),\n    (\"PCA\", PCA(\n        n_components=510,\n        copy=False,\n    )),\n])\n\nX_train = preprocess.fit_transform(X_train)\nX_test = preprocess.transform(X_test)","metadata":{"execution":{"iopub.status.busy":"2023-02-27T18:15:49.300865Z","iopub.execute_input":"2023-02-27T18:15:49.302005Z","iopub.status.idle":"2023-02-27T18:16:02.153268Z","shell.execute_reply.started":"2023-02-27T18:15:49.301963Z","shell.execute_reply":"2023-02-27T18:16:02.151380Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### the benifit Behind using **RandomForestRegressor** \n\n* **use Cases behind chosing this Model**\n\n    1. Non-linear relationships: Random forests can handle non-linear relationships between the features and the target variable, which may be present in the data.\n\n    2. Robust to outliers: Random forests are robust to outliers and noisy data, which can be helpful in situations where the data may contain some degree of noise or anomalies.\n\n    3. Feature selection: Random forests can automatically perform feature selection by evaluating the importance of each feature in predicting the target variable. This can help identify the most relevant features and improve the performance of the model.","metadata":{}},{"cell_type":"code","source":"# Define the estimator and the hyperparameter space\nestimator = RandomForestRegressor(n_estimators=250,\n                                  max_features = 'auto' ,\n                                  max_depth=2,\n                                  min_samples_split = 2,\n                                  min_samples_leaf=3,\n                                  bootstrap=True\n                                 )\n\n# Define the MultiOutputRegressor and the RandomizedSearchCV\nMultiOutputRegress = MultiOutputRegressor(estimator)\n\n# Fit the RandomizedSearchCV on the data\nMultiOutputRegress.fit(X_train, y_train)\n\n\"\"\"\nEvaluating metrices \n\"\"\"# make predictions using the best estimator\ny_pred = MultiOutputRegress.predict(X_test)\n\n# Make predictions on the test data\n# Evaluate the predictions using various metrics\nmse = mean_squared_error(y_test, y_pred)\nmae = mean_absolute_error(y_test, y_pred)\n\n# Print the evaluation metrics\nprint(\"Mean squared error: {:.2f}\".format(mse))\nprint(\"Mean absolute error: {:.2f}\".format(mae))\n","metadata":{"execution":{"iopub.status.busy":"2023-02-27T18:19:41.867166Z","iopub.execute_input":"2023-02-27T18:19:41.867670Z","iopub.status.idle":"2023-02-27T18:20:23.304237Z","shell.execute_reply.started":"2023-02-27T18:19:41.867631Z","shell.execute_reply":"2023-02-27T18:20:23.302711Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del X_train, y_train\ndel X_test, y_test","metadata":{"execution":{"iopub.status.busy":"2023-02-27T18:20:32.507696Z","iopub.execute_input":"2023-02-27T18:20:32.508454Z","iopub.status.idle":"2023-02-27T18:20:32.515925Z","shell.execute_reply.started":"2023-02-27T18:20:32.508409Z","shell.execute_reply":"2023-02-27T18:20:32.514452Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Prepare submission\n\nAn example submission with the correct columns and properly ordered event IDs. The sample submission is provided in the parquet format so it can be read quickly but your final submission must be a csv.\n","metadata":{}},{"cell_type":"code","source":"import gc, time\ngc.collect()\ntime.sleep(9)","metadata":{"execution":{"iopub.status.busy":"2023-02-27T18:20:34.903263Z","iopub.execute_input":"2023-02-27T18:20:34.903739Z","iopub.status.idle":"2023-02-27T18:20:44.175564Z","shell.execute_reply.started":"2023-02-27T18:20:34.903698Z","shell.execute_reply":"2023-02-27T18:20:44.174410Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ssub = pd.read_parquet(os.path.join(PATH_DATASET, \"sample_submission.parquet\"))\n# dtypes={\"azimuth\": np.float16, \"zenith\": np.float16}\nprint(f\"length: {len(ssub)}\")\nssub.set_index(\"event_id\", inplace=True)\ndisplay(ssub.head())\nssub = ssub.apply(np.float16)\ndisplay(ssub.info())","metadata":{"execution":{"iopub.status.busy":"2023-02-27T18:20:44.177296Z","iopub.execute_input":"2023-02-27T18:20:44.178222Z","iopub.status.idle":"2023-02-27T18:20:44.222982Z","shell.execute_reply.started":"2023-02-27T18:20:44.178181Z","shell.execute_reply":"2023-02-27T18:20:44.221525Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# zenith seesms to have Gausina distribution\nssub['zenith'] = meta_data_train['zenith'].mean()","metadata":{"execution":{"iopub.status.busy":"2023-02-27T18:20:51.436811Z","iopub.execute_input":"2023-02-27T18:20:51.437789Z","iopub.status.idle":"2023-02-27T18:20:51.905876Z","shell.execute_reply.started":"2023-02-27T18:20:51.437731Z","shell.execute_reply":"2023-02-27T18:20:51.904388Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ls = glob.glob(os.path.join(PATH_DATASET, \"test\", \"*.parquet\"))\nfor batch_file in ls:\n    print(f\"processing: {batch_file}\")\n    event_ids, data = transform_batch(batch_file)\n    preds = MultiOutputRegress.predict(preprocess.transform(data))\n    del data\n    #print(preds)\n    for eid, (a, z) in zip(event_ids, preds):\n        ssub.at[eid, \"azimuth\"] = a\n        ssub.at[eid, \"zenith\"] = z\n    del event_ids, preds\n    gc.collect()\n    time.sleep(9)","metadata":{"execution":{"iopub.status.busy":"2023-02-27T18:21:09.212771Z","iopub.execute_input":"2023-02-27T18:21:09.213238Z","iopub.status.idle":"2023-02-27T18:21:18.560752Z","shell.execute_reply.started":"2023-02-27T18:21:09.213199Z","shell.execute_reply":"2023-02-27T18:21:18.559679Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ssub.to_csv('submission.csv', index=True)","metadata":{"execution":{"iopub.status.busy":"2023-02-27T18:21:27.460177Z","iopub.execute_input":"2023-02-27T18:21:27.460644Z","iopub.status.idle":"2023-02-27T18:21:27.471847Z","shell.execute_reply.started":"2023-02-27T18:21:27.460603Z","shell.execute_reply":"2023-02-27T18:21:27.470657Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ssubs=pd.read_csv(\"submission.csv\")\nssubs.head(2)","metadata":{"execution":{"iopub.status.busy":"2023-02-27T18:21:29.465669Z","iopub.execute_input":"2023-02-27T18:21:29.466370Z","iopub.status.idle":"2023-02-27T18:21:29.482173Z","shell.execute_reply.started":"2023-02-27T18:21:29.466330Z","shell.execute_reply":"2023-02-27T18:21:29.480364Z"},"trusted":true},"execution_count":null,"outputs":[]}]}