{
  "id": 402864,
  "title": "20th place solution",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/writeups/ice-team-20th-place-solution",
  "author_name": "",
  "post_date": "2023-05-02T21:01:16.527Z",
  "votes": 18,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Team members are junseonglee11 (@junseonglee11), Ayaan Jang(@ayaanjang). <br>\nWe ensembled 6 LSTM models (2 different versions).<br>\nWe modified Robin Smith's and Robert Hatch's notebooks.  </p>\n<h1><strong>Our notebooks:</strong></h1>\n<p>Inference: <a href=\"https://www.kaggle.com/code/ayaanjang/20th-tensorflow-lstm-model-inference-merged\" target=\"_blank\">https://www.kaggle.com/code/ayaanjang/20th-tensorflow-lstm-model-inference-merged</a><br>\nTrain: <a href=\"https://www.kaggle.com/code/junseonglee11/20th-tensorflow-tfrecord-tpu-lstm-line-fit-train\" target=\"_blank\">https://www.kaggle.com/code/junseonglee11/20th-tensorflow-tfrecord-tpu-lstm-line-fit-train</a><br>\nDataset (TFRecord): <a href=\"https://www.kaggle.com/code/junseonglee11/icecube-data-to-tfrecord-v2-1\" target=\"_blank\">https://www.kaggle.com/code/junseonglee11/icecube-data-to-tfrecord-v2-1</a>  </p>\n<h1><strong>References</strong></h1>\n<p>Robin Smith's: notebooks:<br>\n<a href=\"https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-inference\" target=\"_blank\">https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-inference</a><br>\n<a href=\"https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-training-tpu\" target=\"_blank\">https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-training-tpu</a><br>\n<a href=\"https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-data-preprocessor/notebook\" target=\"_blank\">https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-data-preprocessor/notebook</a> I modified his notebook  </p>\n<p>Robert Hatch's notebook<br>\n<a href=\"https://www.kaggle.com/code/roberthatch/lb-1-183-lightning-fast-baseline-with-polars\" target=\"_blank\">https://www.kaggle.com/code/roberthatch/lb-1-183-lightning-fast-baseline-with-polars</a><br>\nIt was crucial to improving our score. Used the results of this notebook as additional inputs in our model.</p>\n<p>Seungmoklee's notebook<br>\n<a href=\"https://www.kaggle.com/code/seungmoklee/lstm-preprocessing-point-picker\" target=\"_blank\">https://www.kaggle.com/code/seungmoklee/lstm-preprocessing-point-picker</a></p>\n<h1><strong>Data preprocessing</strong></h1>\n<p><strong>Our code:</strong> [<a href=\"https://www.kaggle.com/code/junseonglee11/icecube-data-to-tfrecord-v2-1\" target=\"_blank\">https://www.kaggle.com/code/junseonglee11/icecube-data-to-tfrecord-v2-1</a>]</p>\n<ol>\n<li><p>Reference part: Preprocessed data to generate 96 time-series data with 6 features including:</p>\n<ul>\n<li>Sensor signal measurement time</li>\n<li>Sensor signal strength</li>\n<li>Sensor signal quality</li>\n<li>X, Y, Z coordinates of received sensor (3 features)</li></ul></li>\n<li><p>Performed feature engineering on the 6 features to improve the prediction accuracy of the RNN (Residual Neural Network) model (trained on 90 data files instead of the entire dataset, then experimented with feature transformation)</p>\n<ul>\n<li><strong>Using original features: LB 1.015</strong></li>\n<li><strong>Adding Time Difference: LB 1.0128</strong></li>\n<li><strong>Replaced sensor signal measurement time with time interval (time difference between next and current measurement time, Time diff):</strong> <br>\nThe difference in sensor position between current and next time points has different meanings depending on the time interval (for example, <br>\nmoving 10m in 1 second vs. 1m in 10 seconds has a 10-fold speed difference). However, the absolute measurement time of the sensor signal <br>\ncan not reflect this, so it was determined that time interval is a more appropriate input than the measurement time.</li>\n<li><strong>Adding Coordinate Difference: LB 1.0115</strong>    <br>\n<strong>Added three features with the difference in X, Y, Z coordinates between the next and current sensor positions:</strong> Similar to the time <br>\ninterval feature, it was determined that the difference in coordinate values between time points can better reflect the direction information of <br>\nthe neutral particle. However, when replacing the sensor's xyz coordinates with coordinate difference values, the accuracy decreased. This is <br>\nbecause the difference in sensor coordinate values between non-adjacent time points has important information.</li></ul></li>\n<li><p>Convert all the inputs above to TensorFlow TFRecord format to minimize CPU memory usage and accelerate training.</p></li>\n</ol>\n<h1><strong>Model training and inference</strong></h1>\n<p><strong>Our code (model):</strong> <a href=\"https://www.kaggle.com/code/junseonglee11/20th-tensorflow-tfrecord-tpu-lstm-line-fit-train\" target=\"_blank\">https://www.kaggle.com/code/junseonglee11/20th-tensorflow-tfrecord-tpu-lstm-line-fit-train</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7429668%2Feea4665a122c77a5c2e3aa60a38e65b1%2F.png?generation=1683061023145992&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li><strong>Model Overview</strong>: Bidirectional LSTM (long short-term memory) 7 layer+ batch normalization + Line fitting features concatenation + Dense layer</li>\n<li><strong>Activation Function</strong>: GELU(Gaussian Error Linear Unit) activation layer</li>\n<li><strong>Loss Function</strong>: Sparse_categorical_crossentropy loss</li>\n<li><strong>Training Metric</strong>: Accuracy</li>\n<li><strong>Optimizer:</strong> RAdam (Rectified Adam) optimizer</li>\n<li><strong>Ensemble (2 versions of the model)</strong></li>\n<li>Model trained on fold 0 of the dataset divided into 10 folds<ul>\n<li>Train: 594 dataset, Valid: 66 dataset: model trained using only 1 validation dataset</li>\n<li>Train: 659 dataset, Valid: 1 dataset: four models with the lowest Mean Angular Error values are selected from each version of the model (total of 8 models)</li>\n<li>Optimal ensemble weights are determined through cross-validation</li></ul></li>\n</ul>\n<h1><strong>Data postprocessing</strong></h1>\n<p><strong>Our code (Data postprocessing):</strong> <a href=\"https://www.kaggle.com/code/ayaanjang/20th-tensorflow-lstm-model-inference-merged\" target=\"_blank\">https://www.kaggle.com/code/ayaanjang/20th-tensorflow-lstm-model-inference-merged</a><br>\n<strong>Original code:</strong> Power value not squared in the code below.<br>\n     Perform a weighted average of predicted probabilities for each category and the direction of the particle represented by each category to <br>\n     calculate the azimuth and zenith directions of a neutron.<br>\n<strong>Changed code:</strong> Add a square of the power value to the predicted value obtained from the model.<br>\n     It's not optimal to simply derive the results based on the direction for each category and the predicted model probability.<br>\n     Attempt various modifications to the predicted probability using exponential, logarithmic functions, activation functions, etc. to improve the <br>\n     post-processing stage.<br>\n<strong>Improvement:</strong> When the model-predicted category probabilities were squared by 1.35, there was a decrease of about 0.002 in mean angular error.<br>\n→ This gave appropriate additional weight to the category with a high probability in the predicted result.</p>",
  "messages": [
    {
      "id": "2227753",
      "postDate": "04/20/2023 02:08:55",
      "content": "<p>Team members are junseonglee11 (@junseonglee11), Ayaan Jang(@ayaanjang). <br>\nWe ensembled 6 LSTM models (2 different versions).<br>\nWe modified Robin Smith's and Robert Hatch's notebooks.  </p>\n<h1><strong>Our notebooks:</strong></h1>\n<p>Inference: <a href=\"https://www.kaggle.com/code/ayaanjang/20th-tensorflow-lstm-model-inference-merged\" target=\"_blank\">https://www.kaggle.com/code/ayaanjang/20th-tensorflow-lstm-model-inference-merged</a><br>\nTrain: <a href=\"https://www.kaggle.com/code/junseonglee11/20th-tensorflow-tfrecord-tpu-lstm-line-fit-train\" target=\"_blank\">https://www.kaggle.com/code/junseonglee11/20th-tensorflow-tfrecord-tpu-lstm-line-fit-train</a><br>\nDataset (TFRecord): <a href=\"https://www.kaggle.com/code/junseonglee11/icecube-data-to-tfrecord-v2-1\" target=\"_blank\">https://www.kaggle.com/code/junseonglee11/icecube-data-to-tfrecord-v2-1</a>  </p>\n<h1><strong>References</strong></h1>\n<p>Robin Smith's: notebooks:<br>\n<a href=\"https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-inference\" target=\"_blank\">https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-inference</a><br>\n<a href=\"https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-training-tpu\" target=\"_blank\">https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-training-tpu</a><br>\n<a href=\"https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-data-preprocessor/notebook\" target=\"_blank\">https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-data-preprocessor/notebook</a> I modified his notebook  </p>\n<p>Robert Hatch's notebook<br>\n<a href=\"https://www.kaggle.com/code/roberthatch/lb-1-183-lightning-fast-baseline-with-polars\" target=\"_blank\">https://www.kaggle.com/code/roberthatch/lb-1-183-lightning-fast-baseline-with-polars</a><br>\nIt was crucial to improving our score. Used the results of this notebook as additional inputs in our model.</p>\n<p>Seungmoklee's notebook<br>\n<a href=\"https://www.kaggle.com/code/seungmoklee/lstm-preprocessing-point-picker\" target=\"_blank\">https://www.kaggle.com/code/seungmoklee/lstm-preprocessing-point-picker</a></p>\n<h1><strong>Data preprocessing</strong></h1>\n<p><strong>Our code:</strong> [<a href=\"https://www.kaggle.com/code/junseonglee11/icecube-data-to-tfrecord-v2-1\" target=\"_blank\">https://www.kaggle.com/code/junseonglee11/icecube-data-to-tfrecord-v2-1</a>]</p>\n<ol>\n<li><p>Reference part: Preprocessed data to generate 96 time-series data with 6 features including:</p>\n<ul>\n<li>Sensor signal measurement time</li>\n<li>Sensor signal strength</li>\n<li>Sensor signal quality</li>\n<li>X, Y, Z coordinates of received sensor (3 features)</li></ul></li>\n<li><p>Performed feature engineering on the 6 features to improve the prediction accuracy of the RNN (Residual Neural Network) model (trained on 90 data files instead of the entire dataset, then experimented with feature transformation)</p>\n<ul>\n<li><strong>Using original features: LB 1.015</strong></li>\n<li><strong>Adding Time Difference: LB 1.0128</strong></li>\n<li><strong>Replaced sensor signal measurement time with time interval (time difference between next and current measurement time, Time diff):</strong> <br>\nThe difference in sensor position between current and next time points has different meanings depending on the time interval (for example, <br>\nmoving 10m in 1 second vs. 1m in 10 seconds has a 10-fold speed difference). However, the absolute measurement time of the sensor signal <br>\ncan not reflect this, so it was determined that time interval is a more appropriate input than the measurement time.</li>\n<li><strong>Adding Coordinate Difference: LB 1.0115</strong>    <br>\n<strong>Added three features with the difference in X, Y, Z coordinates between the next and current sensor positions:</strong> Similar to the time <br>\ninterval feature, it was determined that the difference in coordinate values between time points can better reflect the direction information of <br>\nthe neutral particle. However, when replacing the sensor's xyz coordinates with coordinate difference values, the accuracy decreased. This is <br>\nbecause the difference in sensor coordinate values between non-adjacent time points has important information.</li></ul></li>\n<li><p>Convert all the inputs above to TensorFlow TFRecord format to minimize CPU memory usage and accelerate training.</p></li>\n</ol>\n<h1><strong>Model training and inference</strong></h1>\n<p><strong>Our code (model):</strong> <a href=\"https://www.kaggle.com/code/junseonglee11/20th-tensorflow-tfrecord-tpu-lstm-line-fit-train\" target=\"_blank\">https://www.kaggle.com/code/junseonglee11/20th-tensorflow-tfrecord-tpu-lstm-line-fit-train</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7429668%2Feea4665a122c77a5c2e3aa60a38e65b1%2F.png?generation=1683061023145992&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li><strong>Model Overview</strong>: Bidirectional LSTM (long short-term memory) 7 layer+ batch normalization + Line fitting features concatenation + Dense layer</li>\n<li><strong>Activation Function</strong>: GELU(Gaussian Error Linear Unit) activation layer</li>\n<li><strong>Loss Function</strong>: Sparse_categorical_crossentropy loss</li>\n<li><strong>Training Metric</strong>: Accuracy</li>\n<li><strong>Optimizer:</strong> RAdam (Rectified Adam) optimizer</li>\n<li><strong>Ensemble (2 versions of the model)</strong></li>\n<li>Model trained on fold 0 of the dataset divided into 10 folds<ul>\n<li>Train: 594 dataset, Valid: 66 dataset: model trained using only 1 validation dataset</li>\n<li>Train: 659 dataset, Valid: 1 dataset: four models with the lowest Mean Angular Error values are selected from each version of the model (total of 8 models)</li>\n<li>Optimal ensemble weights are determined through cross-validation</li></ul></li>\n</ul>\n<h1><strong>Data postprocessing</strong></h1>\n<p><strong>Our code (Data postprocessing):</strong> <a href=\"https://www.kaggle.com/code/ayaanjang/20th-tensorflow-lstm-model-inference-merged\" target=\"_blank\">https://www.kaggle.com/code/ayaanjang/20th-tensorflow-lstm-model-inference-merged</a><br>\n<strong>Original code:</strong> Power value not squared in the code below.<br>\n     Perform a weighted average of predicted probabilities for each category and the direction of the particle represented by each category to <br>\n     calculate the azimuth and zenith directions of a neutron.<br>\n<strong>Changed code:</strong> Add a square of the power value to the predicted value obtained from the model.<br>\n     It's not optimal to simply derive the results based on the direction for each category and the predicted model probability.<br>\n     Attempt various modifications to the predicted probability using exponential, logarithmic functions, activation functions, etc. to improve the <br>\n     post-processing stage.<br>\n<strong>Improvement:</strong> When the model-predicted category probabilities were squared by 1.35, there was a decrease of about 0.002 in mean angular error.<br>\n→ This gave appropriate additional weight to the category with a high probability in the predicted result.</p>",
      "rawMarkdown": "Team members are junseonglee11 (@junseonglee11), Ayaan Jang(@ayaanjang). \nWe ensembled 6 LSTM models (2 different versions).\nWe modified Robin Smith's and Robert Hatch's notebooks.  \n  \n# **Our notebooks:**\nInference: https://www.kaggle.com/code/ayaanjang/20th-tensorflow-lstm-model-inference-merged\nTrain: https://www.kaggle.com/code/junseonglee11/20th-tensorflow-tfrecord-tpu-lstm-line-fit-train\nDataset (TFRecord): https://www.kaggle.com/code/junseonglee11/icecube-data-to-tfrecord-v2-1  \n\n# **References**\nRobin Smith's: notebooks:\nhttps://www.kaggle.com/code/rsmits/tensorflow-lstm-model-inference\nhttps://www.kaggle.com/code/rsmits/tensorflow-lstm-model-training-tpu\nhttps://www.kaggle.com/code/rsmits/tensorflow-lstm-model-data-preprocessor/notebook I modified his notebook  \n  \nRobert Hatch's notebook\nhttps://www.kaggle.com/code/roberthatch/lb-1-183-lightning-fast-baseline-with-polars\nIt was crucial to improving our score. Used the results of this notebook as additional inputs in our model.\n\nSeungmoklee's notebook\nhttps://www.kaggle.com/code/seungmoklee/lstm-preprocessing-point-picker\n\n# **Data preprocessing**\n**Our code:** [https://www.kaggle.com/code/junseonglee11/icecube-data-to-tfrecord-v2-1]\n1. Reference part: Preprocessed data to generate 96 time-series data with 6 features including:\n    - Sensor signal measurement time\n    - Sensor signal strength\n    - Sensor signal quality\n    - X, Y, Z coordinates of received sensor (3 features)\n\n2. Performed feature engineering on the 6 features to improve the prediction accuracy of the RNN (Residual Neural Network) model (trained on 90 data files instead of the entire dataset, then experimented with feature transformation)\n     - **Using original features: LB 1.015**\n     - **Adding Time Difference: LB 1.0128**\n     - **Replaced sensor signal measurement time with time interval (time difference between next and current measurement time, Time diff):** \n       The difference in sensor position between current and next time points has different meanings depending on the time interval (for example, \n       moving 10m in 1 second vs. 1m in 10 seconds has a 10-fold speed difference). However, the absolute measurement time of the sensor signal \n       can not reflect this, so it was determined that time interval is a more appropriate input than the measurement time.\n     - **Adding Coordinate Difference: LB 1.0115**    \n       **Added three features with the difference in X, Y, Z coordinates between the next and current sensor positions:** Similar to the time \n       interval feature, it was determined that the difference in coordinate values between time points can better reflect the direction information of \n       the neutral particle. However, when replacing the sensor's xyz coordinates with coordinate difference values, the accuracy decreased. This is \n       because the difference in sensor coordinate values between non-adjacent time points has important information.\n3. Convert all the inputs above to TensorFlow TFRecord format to minimize CPU memory usage and accelerate training.\n\n\n# **Model training and inference**\n**Our code (model):** https://www.kaggle.com/code/junseonglee11/20th-tensorflow-tfrecord-tpu-lstm-line-fit-train\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7429668%2Feea4665a122c77a5c2e3aa60a38e65b1%2F.png?generation=1683061023145992&alt=media)\n\n- **Model Overview**: Bidirectional LSTM (long short-term memory) 7 layer+ batch normalization + Line fitting features concatenation + Dense layer\n- **Activation Function**: GELU(Gaussian Error Linear Unit) activation layer\n- **Loss Function**: Sparse_categorical_crossentropy loss\n- **Training Metric**: Accuracy\n- **Optimizer:** RAdam (Rectified Adam) optimizer\n- **Ensemble (2 versions of the model)**\n- Model trained on fold 0 of the dataset divided into 10 folds\n     - Train: 594 dataset, Valid: 66 dataset: model trained using only 1 validation dataset\n     - Train: 659 dataset, Valid: 1 dataset: four models with the lowest Mean Angular Error values are selected from each version of the model (total of 8 models)\n     -  Optimal ensemble weights are determined through cross-validation\n\n\n# **Data postprocessing**\n **Our code (Data postprocessing):** [https://www.kaggle.com/code/ayaanjang/20th-tensorflow-lstm-model-inference-merged](https://www.kaggle.com/code/ayaanjang/20th-tensorflow-lstm-model-inference-merged)\n**Original code:** Power value not squared in the code below.\n     Perform a weighted average of predicted probabilities for each category and the direction of the particle represented by each category to \n     calculate the azimuth and zenith directions of a neutron.\n**Changed code:** Add a square of the power value to the predicted value obtained from the model.\n     It's not optimal to simply derive the results based on the direction for each category and the predicted model probability.\n     Attempt various modifications to the predicted probability using exponential, logarithmic functions, activation functions, etc. to improve the \n     post-processing stage.\n**Improvement:** When the model-predicted category probabilities were squared by 1.35, there was a decrease of about 0.002 in mean angular error.\n→ This gave appropriate additional weight to the category with a high probability in the predicted result.",
      "votes": null
    },
    {
      "id": "2228452",
      "postDate": "04/20/2023 15:00:35",
      "content": "<p><a href=\"https://www.kaggle.com/junseonglee11\" target=\"_blank\">@junseonglee11</a> nice to see, that you pushed LSTMs that high. Could you explain a bit more this item: <em>Additional inputs and some preprocessing</em>? I tried many experiments with LSTMs, but the best ensemble was around 0.998.</p>",
      "rawMarkdown": "junseonglee11 nice to see, that you pushed LSTMs that high. Could you explain a bit more this item: *Additional inputs and some preprocessing*? I tried many experiments with LSTMs, but the best ensemble was around 0.998.",
      "votes": null
    },
    {
      "id": "2228849",
      "postDate": "04/20/2023 21:43:18",
      "content": "<p><a href=\"https://www.kaggle.com/manwithaflower\" target=\"_blank\">@manwithaflower</a> additional preprocessing is as follows.</p>\n<ol>\n<li>Used time interval (next time point - current time point) instead of time point (becuase for RNN, I felt that interval is more natural.)</li>\n<li>Added difference of sensor position of next and current time point (because I thought that it may contain direction information)</li>\n<li>Added line fit result of Robert Hatch's notebook and concatenated it to the layer after the last LSTM layer.</li>\n</ol>\n<p>Also, stacking 7 LSTM layers and using all data with TFRecord format were crucial.</p>\n<p>Thank you!</p>",
      "rawMarkdown": "manwithaflower additional preprocessing is as follows.\n1. Used time interval (next time point - current time point) instead of time point (becuase for RNN, I felt that interval is more natural.)\n2. Added difference of sensor position of next and current time point (because I thought that it may contain direction information)\n3. Added line fit result of Robert Hatch's notebook and concatenated it to the layer after the last LSTM layer.\n\nAlso, stacking 7 LSTM layers and using all data with TFRecord format were crucial.\n\nThank you!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2228452,
      "author_name": "manwithaflower",
      "author_url": "",
      "post_date": "04/20/2023 15:00:35",
      "content": "<p><a href=\"https://www.kaggle.com/junseonglee11\" target=\"_blank\">@junseonglee11</a> nice to see, that you pushed LSTMs that high. Could you explain a bit more this item: <em>Additional inputs and some preprocessing</em>? I tried many experiments with LSTMs, but the best ensemble was around 0.998.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2228849,
          "author_name": "junseonglee11",
          "author_url": "",
          "post_date": "04/20/2023 21:43:18",
          "content": "<p><a href=\"https://www.kaggle.com/manwithaflower\" target=\"_blank\">@manwithaflower</a> additional preprocessing is as follows.</p>\n<ol>\n<li>Used time interval (next time point - current time point) instead of time point (becuase for RNN, I felt that interval is more natural.)</li>\n<li>Added difference of sensor position of next and current time point (because I thought that it may contain direction information)</li>\n<li>Added line fit result of Robert Hatch's notebook and concatenated it to the layer after the last LSTM layer.</li>\n</ol>\n<p>Also, stacking 7 LSTM layers and using all data with TFRecord format were crucial.</p>\n<p>Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2227753": "Team members are junseonglee11 (@junseonglee11), Ayaan Jang(@ayaanjang). \nWe ensembled 6 LSTM models (2 different versions).\nWe modified Robin Smith's and Robert Hatch's notebooks.  \n  \n# **Our notebooks:**\nInference: https://www.kaggle.com/code/ayaanjang/20th-tensorflow-lstm-model-inference-merged\nTrain: https://www.kaggle.com/code/junseonglee11/20th-tensorflow-tfrecord-tpu-lstm-line-fit-train\nDataset (TFRecord): https://www.kaggle.com/code/junseonglee11/icecube-data-to-tfrecord-v2-1  \n\n# **References**\nRobin Smith's: notebooks:\nhttps://www.kaggle.com/code/rsmits/tensorflow-lstm-model-inference\nhttps://www.kaggle.com/code/rsmits/tensorflow-lstm-model-training-tpu\nhttps://www.kaggle.com/code/rsmits/tensorflow-lstm-model-data-preprocessor/notebook I modified his notebook  \n  \nRobert Hatch's notebook\nhttps://www.kaggle.com/code/roberthatch/lb-1-183-lightning-fast-baseline-with-polars\nIt was crucial to improving our score. Used the results of this notebook as additional inputs in our model.\n\nSeungmoklee's notebook\nhttps://www.kaggle.com/code/seungmoklee/lstm-preprocessing-point-picker\n\n# **Data preprocessing**\n**Our code:** [https://www.kaggle.com/code/junseonglee11/icecube-data-to-tfrecord-v2-1]\n1. Reference part: Preprocessed data to generate 96 time-series data with 6 features including:\n    - Sensor signal measurement time\n    - Sensor signal strength\n    - Sensor signal quality\n    - X, Y, Z coordinates of received sensor (3 features)\n\n2. Performed feature engineering on the 6 features to improve the prediction accuracy of the RNN (Residual Neural Network) model (trained on 90 data files instead of the entire dataset, then experimented with feature transformation)\n     - **Using original features: LB 1.015**\n     - **Adding Time Difference: LB 1.0128**\n     - **Replaced sensor signal measurement time with time interval (time difference between next and current measurement time, Time diff):** \n       The difference in sensor position between current and next time points has different meanings depending on the time interval (for example, \n       moving 10m in 1 second vs. 1m in 10 seconds has a 10-fold speed difference). However, the absolute measurement time of the sensor signal \n       can not reflect this, so it was determined that time interval is a more appropriate input than the measurement time.\n     - **Adding Coordinate Difference: LB 1.0115**    \n       **Added three features with the difference in X, Y, Z coordinates between the next and current sensor positions:** Similar to the time \n       interval feature, it was determined that the difference in coordinate values between time points can better reflect the direction information of \n       the neutral particle. However, when replacing the sensor's xyz coordinates with coordinate difference values, the accuracy decreased. This is \n       because the difference in sensor coordinate values between non-adjacent time points has important information.\n3. Convert all the inputs above to TensorFlow TFRecord format to minimize CPU memory usage and accelerate training.\n\n\n# **Model training and inference**\n**Our code (model):** https://www.kaggle.com/code/junseonglee11/20th-tensorflow-tfrecord-tpu-lstm-line-fit-train\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7429668%2Feea4665a122c77a5c2e3aa60a38e65b1%2F.png?generation=1683061023145992&alt=media)\n\n- **Model Overview**: Bidirectional LSTM (long short-term memory) 7 layer+ batch normalization + Line fitting features concatenation + Dense layer\n- **Activation Function**: GELU(Gaussian Error Linear Unit) activation layer\n- **Loss Function**: Sparse_categorical_crossentropy loss\n- **Training Metric**: Accuracy\n- **Optimizer:** RAdam (Rectified Adam) optimizer\n- **Ensemble (2 versions of the model)**\n- Model trained on fold 0 of the dataset divided into 10 folds\n     - Train: 594 dataset, Valid: 66 dataset: model trained using only 1 validation dataset\n     - Train: 659 dataset, Valid: 1 dataset: four models with the lowest Mean Angular Error values are selected from each version of the model (total of 8 models)\n     -  Optimal ensemble weights are determined through cross-validation\n\n\n# **Data postprocessing**\n **Our code (Data postprocessing):** [https://www.kaggle.com/code/ayaanjang/20th-tensorflow-lstm-model-inference-merged](https://www.kaggle.com/code/ayaanjang/20th-tensorflow-lstm-model-inference-merged)\n**Original code:** Power value not squared in the code below.\n     Perform a weighted average of predicted probabilities for each category and the direction of the particle represented by each category to \n     calculate the azimuth and zenith directions of a neutron.\n**Changed code:** Add a square of the power value to the predicted value obtained from the model.\n     It's not optimal to simply derive the results based on the direction for each category and the predicted model probability.\n     Attempt various modifications to the predicted probability using exponential, logarithmic functions, activation functions, etc. to improve the \n     post-processing stage.\n**Improvement:** When the model-predicted category probabilities were squared by 1.35, there was a decrease of about 0.002 in mean angular error.\n→ This gave appropriate additional weight to the category with a high probability in the predicted result.",
    "2228452": "junseonglee11 nice to see, that you pushed LSTMs that high. Could you explain a bit more this item: *Additional inputs and some preprocessing*? I tried many experiments with LSTMs, but the best ensemble was around 0.998.",
    "2228849": "manwithaflower additional preprocessing is as follows.\n1. Used time interval (next time point - current time point) instead of time point (becuase for RNN, I felt that interval is more natural.)\n2. Added difference of sensor position of next and current time point (because I thought that it may contain direction information)\n3. Added line fit result of Robert Hatch's notebook and concatenated it to the layer after the last LSTM layer.\n\nAlso, stacking 7 LSTM layers and using all data with TFRecord format were crucial.\n\nThank you!"
  },
  "source": "meta"
}