{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### My ideal:\n* In this model, I tried to use sensor data to predict direction neutrinos. Sensor data may be very important for suggest direction. As example, if a neutrino beam is come from west north to north, it will never activate a sensor in south pole. So by filtering sensor data, we can suggest some directions for neutrino beams.\n### Sensor embedding data\n* Sensor data is embedded by using PCA. But due to huge usage memory, I embed it in a seperate notebook with only CPU kernel: https://www.kaggle.com/code/astrung/make-pca-data/notebook . Please upvote it if you want to use this data\n### Model\n*  In this notebook, we will use embedding sensor data with catboost and optuna. Catboost is regression model, and optuna is used for auto tunning. For better speed, I used GPU too.\n### Result:\n* As you see in tunning log, our model accuracy is nearly same with every model parameters. It suggest that may be our embedding data isn't good enough to suggest differences between direction of beams.\n* I am thinking about some new approach: using deep learning/BERT embedding, or direct applying a linear regression for each event. \n* If you have another embedding, please suggest with me in this topic: https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/381078","metadata":{}},{"cell_type":"code","source":"%matplotlib inline\n\nimport os\nimport glob\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport gc\n\nPATH_DATASET = \"/kaggle/input/icecube-neutrinos-in-deep-ice\"","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tqdm.auto import tqdm\n\ndef transform_batch(path_batch_parquet, nb_sensors=5160):\n    df = pd.read_parquet(path_batch_parquet)\n    # df[df['auxiliary']]['charge'] /= 10.\n    data = np.zeros((len(df.index.unique()), nb_sensors), dtype=np.float16)\n    event_ids = []\n    for i, (idx, dfg) in tqdm(enumerate(df.groupby(level=0))):\n        event_ids.append(idx)\n        data[i, dfg['sensor_id']] = dfg['charge']\n    # df = pd.DataFrame(data, index=event_ids, columns=[f\"sid-{i}\" for i in range(nb_sensors)])\n    return event_ids, data","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Data preprocessing","metadata":{}},{"cell_type":"markdown","source":"I load my PCA embedded data from another notebook result. If you want to check about how to embed with PCA, please check this notebook: https://www.kaggle.com/code/astrung/make-pca-data/notebook","metadata":{}},{"cell_type":"code","source":"X_train = np.load('/kaggle/input/make-pca-data/x_train.npy')\nX_test = np.load('/kaggle/input/make-pca-data/x_test.npy')\ny_train = np.load('/kaggle/input/make-pca-data/y_train.npy')\ny_test = np.load('/kaggle/input/make-pca-data/y_test.npy')\nX_train.shape, y_train.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# from sklearn.decomposition import PCA, TruncatedSVD, IncrementalPCA\n# from sklearn.pipeline import Pipeline\n# from sklearn.preprocessing import StandardScaler\n# from xgboost import XGBRegressor\n\n# preprocess = Pipeline([\n#     ('scaler', StandardScaler()),\n#     # https://scikit-learn.org/stable/modules/generated/sklearn.decomposition.PCA.html\n#     (\"PCA\", PCA(\n#         n_components=510,\n#         copy=False,\n#     )),\n# #     (\"PCA\", IncrementalPCA(\n# #         n_components=128,\n# #         copy=False,\n# #         batch_size=128\n# #     )),\n# #     (\"PCA\", TruncatedSVD(n_components=510, algorithm='arpack'))\n# ])\n\n# X_train = preprocess.fit_transform(X_train)\n# X_test = preprocess.transform(X_test)\n# X_train.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_train, y_train[:, 0]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Regression baseline\n\n","metadata":{}},{"cell_type":"markdown","source":"Now we use optuna for tuning, but accuracy doesn't chage. May be our embedding isn't good :(","metadata":{}},{"cell_type":"code","source":"from catboost import CatBoost, CatBoostRegressor\nimport sklearn\nimport optuna","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"best_param = {\n#  'task_type': 'GPU',\n 'verbose': False,\n 'loss_function': 'RMSE',\n#  'subsample': 1,\n 'iterations': 90,\n 'learning_rate': 0.00516987251960874,\n 'depth': 5,\n 'l2_leaf_reg': 0.7160558941109652,\n 'grow_policy': 'SymmetricTree',\n 'min_data_in_leaf': 77}\nmodel = CatBoostRegressor(**best_param)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.fit(X_train, y_train[:, 0])\npred_labels = model.predict(X_test)\npred_labels","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sklearn.metrics.mean_absolute_error(y_test[:, 0], pred_labels)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":" def objective_azimuth(trial):\n    # 2. Suggest values of the hyperparameters using a trial object.\n    param = {\n#         'task_type': 'GPU',\n        'verbose': False,\n#         'auto_class_weights':trial.suggest_categorical('auto_class_weights', [None, 'SqrtBalanced', 'Balanced']),\n#         'boosting-type': 'Plain',\n#         \"used_ram_limit\": \"3gb\",\n#         'per_float_feature_quantization': ['0:border_count=1024', '1:border_count=1024'],\n        'iterations': trial.suggest_int('iterations', 10, 2000, step=10),\n        \"learning_rate\": trial.suggest_float(\"learning_rate\", 1e-5, 1e0),\n        \"depth\": trial.suggest_int(\"depth\", 2, 12),\n        'l2_leaf_reg': trial.suggest_float(\"l2_leaf_reg\", 1e-2, 1e2, log=True),\n        \"grow_policy\": trial.suggest_categorical(\"grow_policy\", [\"SymmetricTree\", \"Depthwise\", \"Lossguide\"]),\n#         \"has_time\": trial.suggest_categorical(\"has_time\", [True, False]),\n#         \"bagging_temperature\": trial.suggest_int(\"bagging_temperature\", 1, 1e3),\n        \"border_count\": trial.suggest_int(\"border_count\", 1, 254),\n        \"min_data_in_leaf\": trial.suggest_int(\"min_data_in_leaf\", 1, 199, step=2),\n#         'one_hot_max_size':trial.suggest_int(\"one_hot_max_size\", 100, 250, step=10),\n        'loss_function': trial.suggest_categorical(\"loss_function\", ['RMSE'])\n#         'loss_function': trial.suggest_categorical(\"loss_function\", ['RMSE', 'MAPE', 'LogLinQuantile', 'Quantile']),\n#         'loss_function': 'Quantile:alpha='+str(trial.suggest_float(\"alpha\", 1e-5, 1)),\n#         \"colsample_bylevel\": trial.suggest_float(\"colsample_bylevel\", 0.01, 0.1, log=True)\n#         \"max_leaves\": trial.suggest_int(\"max_leaves\", 1, 100, step=2)\n    }\n    print(param)\n    model = CatBoostRegressor(**param)\n    model.fit(X_train, y_train[:, 0])\n    pred_labels = model.predict(X_test)\n#     pred_labels = np.mean(pred_labels, axis=1)\n    accuracy = sklearn.metrics.mean_absolute_error(y_test[:, 0], pred_labels)\n    return accuracy","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"study = optuna.create_study(direction=\"minimize\")\nstudy.optimize(objective_azimuth, n_trials=10, gc_after_trial=True)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"trial = study.best_trial\ntrial.params","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"params = {\n#     'task_type': 'GPU',\n    'verbose': False,\n}\nparams.update(**trial.params)\n# params.update(**best_param)\nparams","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model_azimuth = CatBoostRegressor(**params)\nprint(model_azimuth.get_params())\nmodel_azimuth.fit(X_train, y_train[:, 0])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":" def objective_zenith(trial):\n    # 2. Suggest values of the hyperparameters using a trial object.\n    param = {\n#         'task_type': 'GPU',\n        'verbose': False,\n#         'auto_class_weights':trial.suggest_categorical('auto_class_weights', [None, 'SqrtBalanced', 'Balanced']),\n#         'boosting-type': 'Plain',\n#         \"used_ram_limit\": \"3gb\",\n#         'per_float_feature_quantization': ['0:border_count=1024', '1:border_count=1024'],\n        'iterations': trial.suggest_int('iterations', 10, 2000, step=10),\n        \"learning_rate\": trial.suggest_float(\"learning_rate\", 1e-5, 1e0),\n        \"depth\": trial.suggest_int(\"depth\", 2, 12),\n        'l2_leaf_reg': trial.suggest_float(\"l2_leaf_reg\", 1e-2, 1e2, log=True),\n        \"grow_policy\": trial.suggest_categorical(\"grow_policy\", [\"SymmetricTree\", \"Depthwise\", \"Lossguide\"]),\n#         \"has_time\": trial.suggest_categorical(\"has_time\", [True, False]),\n#         \"bagging_temperature\": trial.suggest_int(\"bagging_temperature\", 1, 1e3),\n        \"border_count\": trial.suggest_int(\"border_count\", 1, 254),\n        \"min_data_in_leaf\": trial.suggest_int(\"min_data_in_leaf\", 1, 199, step=2),\n#         'one_hot_max_size':trial.suggest_int(\"one_hot_max_size\", 100, 250, step=10),\n        'loss_function': trial.suggest_categorical(\"loss_function\", ['RMSE'])\n#         'loss_function': trial.suggest_categorical(\"loss_function\", ['RMSE', 'MAPE', 'LogLinQuantile', 'Quantile']),\n#         'loss_function': 'Quantile:alpha='+str(trial.suggest_float(\"alpha\", 1e-5, 1)),\n#         \"colsample_bylevel\": trial.suggest_float(\"colsample_bylevel\", 0.01, 0.1, log=True)\n#         \"max_leaves\": trial.suggest_int(\"max_leaves\", 1, 100, step=2)\n    }\n    print(param)\n    model = CatBoostRegressor(**param)\n    model.fit(X_train, y_train[:, 1])\n    pred_labels = model.predict(X_test)\n#     pred_labels = np.mean(pred_labels, axis=1)\n    accuracy = sklearn.metrics.mean_absolute_error(y_test[:, 1], pred_labels)\n    return accuracy","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"study = optuna.create_study(direction=\"minimize\")\nstudy.optimize(objective_zenith, n_trials=10, gc_after_trial=True)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"trial = study.best_trial\nprint(trial.params)\nparams = {\n#     'task_type': 'GPU',\n    'verbose': False,\n}\nparams.update(**trial.params)\n# params.update(**best_param)\nparams","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model_zenith = CatBoostRegressor(**params)\nprint(model_zenith.get_params())\nmodel_zenith.fit(X_train, y_train[:, 1])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Prepare submission\n\nAn example submission with the correct columns and properly ordered event IDs. The sample submission is provided in the parquet format so it can be read quickly but your final submission must be a csv.","metadata":{}},{"cell_type":"code","source":"ssub = pd.read_parquet(os.path.join(PATH_DATASET, \"sample_submission.parquet\"))\n# dtypes={\"azimuth\": np.float16, \"zenith\": np.float16}\nprint(f\"length: {len(ssub)}\")\nssub.set_index(\"event_id\", inplace=True)\ndisplay(ssub.head())\nssub = ssub.apply(np.float16)\ndisplay(ssub.info())","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Fill defaults","metadata":{}},{"cell_type":"code","source":"import pickle\nwith open(r\"/kaggle/input/make-pca-data/preprocessor\", \"rb\") as input_file:\n    preprocess = pickle.load(input_file)\npreprocess","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Predictions / inference","metadata":{}},{"cell_type":"code","source":"ls = glob.glob(os.path.join(PATH_DATASET, \"test\", \"*.parquet\"))\n\nfor batch_file in ls:\n    print(f\"processing: {batch_file}\")\n    event_ids, data = transform_batch(batch_file)\n    data = preprocess.transform(data)\n    azimuth_preds = model_azimuth.predict(data)\n    zenith_preds = model_zenith.predict(data)\n    del data\n    #print(preds)\n    for event_id, a, z in zip(event_ids, azimuth_preds, zenith_preds):\n        ssub.at[event_id, \"azimuth\"] = a\n        ssub.at[event_id, \"zenith\"] = z\n    del event_ids, zenith_preds, azimuth_preds\n    gc.collect()\n#     time.sleep(9)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ssub","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ssub.to_csv('submission.csv', index=True)\n\n!head submission.csv","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}