{
  "id": 405563,
  "title": "13th Place Solution",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/writeups/sinpcw-13th-place-solution",
  "author_name": "",
  "post_date": "2023-05-03T09:31:06.363Z",
  "votes": 16,
  "comment_count": 2,
  "views": 0,
  "content": "<p>A thank you to the host, and well done to all the participants. I'm frustrated that I couldn't stay in the gold tier, but I gained valuable learning experiences. I'd like to take this opportunity to express my gratitude. Although it's late, I will share my solution. My English isn't so good so feel free to ask me if there is anything unclear.</p>\n<p>My solution is:</p>\n<ul>\n<li>Noise reduction from training data</li>\n<li>Customized GraphNet-based models</li>\n</ul>\n<p>I tried training LSTM and Transformers, but it didn't go well, so I abandoned them early, which led to my failure.</p>\n<p>I used an ensemble of eight GraphNet-based models on final submit. A single graph model get Public 0.982 and Private 0.984 scores. Therefore, there may not have been much of an ensemble effect.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2391815%2Fc25aef115dd90773cb76ae7fc0ae296c%2Flatesub.png?generation=1682672834486253&amp;alt=media\" alt=\"late sub\"></p>\n<p><strong>Noise reduction from training data</strong><br>\nI consider when the data had few observed signals within an event, it couldn't fit well and led to overfitting. So, I implemented the following cycle:<br>\n1) Infer to training data and exclude training data that do not predict well. Specifically, exclude data with MAE &gt; 1.<br>\n2) Train on the dataset, and if the validation score improves, use that model to infer the training data again.<br>\n3) Construct the data set in the same manner as step 1, exclude the training data that cannot be fitted, and train again.</p>\n<p>The above process was repeated two or three times.</p>\n<p>This idea is based on my past experience with overfitting with noisy labels, as in the PANDA competition. I think this competition deals with physical phenomena, prediction is relatively easy given ideal observational data. Therefore, I considered a model trained include data like noise would not correctly predict difficult events on test data, too.  In fact, the model trained with the noisy data was able to improve the CV. Public LB also improved as did CV, so I trusts this method.</p>\n<p><strong>Customized GraphNet-based models</strong><br>\nI created models based on provided by the host, I also created models that include time variables in the information added to the KNN. And, changed the number of KNN neighborhoods (8 to 32) and pooling layers(replacing pooling to attention). I also tried other things such as changing the activation function, but the zero gradient function in the negative region does not seem to be a good. PReLU was the best, in my case.</p>\n<p><strong>My Failure Points</strong></p>\n<ul>\n<li>Other information<br>\nTried to include information on QE, etc., but could not pursue it very deeply, especially since GNN did not seem to improve.</li>\n<li>Other approach<br>\nI could not train LSTM and Transformer well. It is possible that the inputs and settings were incorrect, but the biggest failure was to terminate this approach prematurely.</li>\n<li>post-process<br>\nAttempted post-processing in the classification task to deal with cases where the exact opposite direction was obtained, but it did not work. </li>\n<li>Interpolation<br>\nSince the accuracy was poor for events with few observed signals, I looked for ways to improve the accuracy by interpolating signals at intermediate times and intermediate locations, but this did not work.</li>\n<li>Imaging Idea<br>\nTried to create an ensemble factor by predicting the position and orientation of the sensor as a video/imaging task, but it did not work because the number of events was too large (storage was disk full). I should have focused on something else.</li>\n</ul>\n<p><strong>Others</strong></p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/sinpcw/icecube-submit\" target=\"_blank\">inference notebook</a> (version 111 is best LB.)</li>\n</ul>",
  "messages": [
    {
      "id": "2238236",
      "postDate": "04/28/2023 10:32:40",
      "content": "<p>A thank you to the host, and well done to all the participants. I'm frustrated that I couldn't stay in the gold tier, but I gained valuable learning experiences. I'd like to take this opportunity to express my gratitude. Although it's late, I will share my solution. My English isn't so good so feel free to ask me if there is anything unclear.</p>\n<p>My solution is:</p>\n<ul>\n<li>Noise reduction from training data</li>\n<li>Customized GraphNet-based models</li>\n</ul>\n<p>I tried training LSTM and Transformers, but it didn't go well, so I abandoned them early, which led to my failure.</p>\n<p>I used an ensemble of eight GraphNet-based models on final submit. A single graph model get Public 0.982 and Private 0.984 scores. Therefore, there may not have been much of an ensemble effect.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2391815%2Fc25aef115dd90773cb76ae7fc0ae296c%2Flatesub.png?generation=1682672834486253&amp;alt=media\" alt=\"late sub\"></p>\n<p><strong>Noise reduction from training data</strong><br>\nI consider when the data had few observed signals within an event, it couldn't fit well and led to overfitting. So, I implemented the following cycle:<br>\n1) Infer to training data and exclude training data that do not predict well. Specifically, exclude data with MAE &gt; 1.<br>\n2) Train on the dataset, and if the validation score improves, use that model to infer the training data again.<br>\n3) Construct the data set in the same manner as step 1, exclude the training data that cannot be fitted, and train again.</p>\n<p>The above process was repeated two or three times.</p>\n<p>This idea is based on my past experience with overfitting with noisy labels, as in the PANDA competition. I think this competition deals with physical phenomena, prediction is relatively easy given ideal observational data. Therefore, I considered a model trained include data like noise would not correctly predict difficult events on test data, too.  In fact, the model trained with the noisy data was able to improve the CV. Public LB also improved as did CV, so I trusts this method.</p>\n<p><strong>Customized GraphNet-based models</strong><br>\nI created models based on provided by the host, I also created models that include time variables in the information added to the KNN. And, changed the number of KNN neighborhoods (8 to 32) and pooling layers(replacing pooling to attention). I also tried other things such as changing the activation function, but the zero gradient function in the negative region does not seem to be a good. PReLU was the best, in my case.</p>\n<p><strong>My Failure Points</strong></p>\n<ul>\n<li>Other information<br>\nTried to include information on QE, etc., but could not pursue it very deeply, especially since GNN did not seem to improve.</li>\n<li>Other approach<br>\nI could not train LSTM and Transformer well. It is possible that the inputs and settings were incorrect, but the biggest failure was to terminate this approach prematurely.</li>\n<li>post-process<br>\nAttempted post-processing in the classification task to deal with cases where the exact opposite direction was obtained, but it did not work. </li>\n<li>Interpolation<br>\nSince the accuracy was poor for events with few observed signals, I looked for ways to improve the accuracy by interpolating signals at intermediate times and intermediate locations, but this did not work.</li>\n<li>Imaging Idea<br>\nTried to create an ensemble factor by predicting the position and orientation of the sensor as a video/imaging task, but it did not work because the number of events was too large (storage was disk full). I should have focused on something else.</li>\n</ul>\n<p><strong>Others</strong></p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/sinpcw/icecube-submit\" target=\"_blank\">inference notebook</a> (version 111 is best LB.)</li>\n</ul>",
      "rawMarkdown": "A thank you to the host, and well done to all the participants. I'm frustrated that I couldn't stay in the gold tier, but I gained valuable learning experiences. I'd like to take this opportunity to express my gratitude. Although it's late, I will share my solution. My English isn't so good so feel free to ask me if there is anything unclear.\n\nMy solution is:\n- Noise reduction from training data\n- Customized GraphNet-based models\n\nI tried training LSTM and Transformers, but it didn't go well, so I abandoned them early, which led to my failure.\n\nI used an ensemble of eight GraphNet-based models on final submit. A single graph model get Public 0.982 and Private 0.984 scores. Therefore, there may not have been much of an ensemble effect.\n![late sub](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2391815%2Fc25aef115dd90773cb76ae7fc0ae296c%2Flatesub.png?generation=1682672834486253&alt=media)\n\n\n**Noise reduction from training data**\nI consider when the data had few observed signals within an event, it couldn't fit well and led to overfitting. So, I implemented the following cycle:\n1) Infer to training data and exclude training data that do not predict well. Specifically, exclude data with MAE > 1.\n2) Train on the dataset, and if the validation score improves, use that model to infer the training data again.\n3) Construct the data set in the same manner as step 1, exclude the training data that cannot be fitted, and train again.\n\nThe above process was repeated two or three times.\n\nThis idea is based on my past experience with overfitting with noisy labels, as in the PANDA competition. I think this competition deals with physical phenomena, prediction is relatively easy given ideal observational data. Therefore, I considered a model trained include data like noise would not correctly predict difficult events on test data, too.  In fact, the model trained with the noisy data was able to improve the CV. Public LB also improved as did CV, so I trusts this method.\n\n**Customized GraphNet-based models**\nI created models based on provided by the host, I also created models that include time variables in the information added to the KNN. And, changed the number of KNN neighborhoods (8 to 32) and pooling layers(replacing pooling to attention). I also tried other things such as changing the activation function, but the zero gradient function in the negative region does not seem to be a good. PReLU was the best, in my case.\n\n**My Failure Points**\n- Other information\nTried to include information on QE, etc., but could not pursue it very deeply, especially since GNN did not seem to improve.\n- Other approach\nI could not train LSTM and Transformer well. It is possible that the inputs and settings were incorrect, but the biggest failure was to terminate this approach prematurely.\n- post-process\nAttempted post-processing in the classification task to deal with cases where the exact opposite direction was obtained, but it did not work. \n- Interpolation\nSince the accuracy was poor for events with few observed signals, I looked for ways to improve the accuracy by interpolating signals at intermediate times and intermediate locations, but this did not work.\n- Imaging Idea\nTried to create an ensemble factor by predicting the position and orientation of the sensor as a video/imaging task, but it did not work because the number of events was too large (storage was disk full). I should have focused on something else.\n\n**Others**\n- [inference notebook](https://www.kaggle.com/code/sinpcw/icecube-submit) (version 111 is best LB.)",
      "votes": null
    },
    {
      "id": "2238320",
      "postDate": "04/28/2023 12:16:17",
      "content": "<p><a href=\"https://www.kaggle.com/sinpcw\" target=\"_blank\">@sinpcw</a> thank you for the topic. Very interesting reading. And very good that you write about failure points. It is good information for beginners.</p>",
      "rawMarkdown": "sinpcw thank you for the topic. Very interesting reading. And very good that you write about failure points. It is good information for beginners.",
      "votes": null
    },
    {
      "id": "2239181",
      "postDate": "04/29/2023 08:30:49",
      "content": "<p>Thanks for sharing this topic. This will be helpful for beginners who are just starting out.</p>",
      "rawMarkdown": "Thanks for sharing this topic. This will be helpful for beginners who are just starting out.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2238320,
      "author_name": "serangu",
      "author_url": "",
      "post_date": "04/28/2023 12:16:17",
      "content": "<p><a href=\"https://www.kaggle.com/sinpcw\" target=\"_blank\">@sinpcw</a> thank you for the topic. Very interesting reading. And very good that you write about failure points. It is good information for beginners.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2239181,
      "author_name": "ericka42",
      "author_url": "",
      "post_date": "04/29/2023 08:30:49",
      "content": "<p>Thanks for sharing this topic. This will be helpful for beginners who are just starting out.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2238236": "A thank you to the host, and well done to all the participants. I'm frustrated that I couldn't stay in the gold tier, but I gained valuable learning experiences. I'd like to take this opportunity to express my gratitude. Although it's late, I will share my solution. My English isn't so good so feel free to ask me if there is anything unclear.\n\nMy solution is:\n- Noise reduction from training data\n- Customized GraphNet-based models\n\nI tried training LSTM and Transformers, but it didn't go well, so I abandoned them early, which led to my failure.\n\nI used an ensemble of eight GraphNet-based models on final submit. A single graph model get Public 0.982 and Private 0.984 scores. Therefore, there may not have been much of an ensemble effect.\n![late sub](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2391815%2Fc25aef115dd90773cb76ae7fc0ae296c%2Flatesub.png?generation=1682672834486253&alt=media)\n\n\n**Noise reduction from training data**\nI consider when the data had few observed signals within an event, it couldn't fit well and led to overfitting. So, I implemented the following cycle:\n1) Infer to training data and exclude training data that do not predict well. Specifically, exclude data with MAE > 1.\n2) Train on the dataset, and if the validation score improves, use that model to infer the training data again.\n3) Construct the data set in the same manner as step 1, exclude the training data that cannot be fitted, and train again.\n\nThe above process was repeated two or three times.\n\nThis idea is based on my past experience with overfitting with noisy labels, as in the PANDA competition. I think this competition deals with physical phenomena, prediction is relatively easy given ideal observational data. Therefore, I considered a model trained include data like noise would not correctly predict difficult events on test data, too.  In fact, the model trained with the noisy data was able to improve the CV. Public LB also improved as did CV, so I trusts this method.\n\n**Customized GraphNet-based models**\nI created models based on provided by the host, I also created models that include time variables in the information added to the KNN. And, changed the number of KNN neighborhoods (8 to 32) and pooling layers(replacing pooling to attention). I also tried other things such as changing the activation function, but the zero gradient function in the negative region does not seem to be a good. PReLU was the best, in my case.\n\n**My Failure Points**\n- Other information\nTried to include information on QE, etc., but could not pursue it very deeply, especially since GNN did not seem to improve.\n- Other approach\nI could not train LSTM and Transformer well. It is possible that the inputs and settings were incorrect, but the biggest failure was to terminate this approach prematurely.\n- post-process\nAttempted post-processing in the classification task to deal with cases where the exact opposite direction was obtained, but it did not work. \n- Interpolation\nSince the accuracy was poor for events with few observed signals, I looked for ways to improve the accuracy by interpolating signals at intermediate times and intermediate locations, but this did not work.\n- Imaging Idea\nTried to create an ensemble factor by predicting the position and orientation of the sensor as a video/imaging task, but it did not work because the number of events was too large (storage was disk full). I should have focused on something else.\n\n**Others**\n- [inference notebook](https://www.kaggle.com/code/sinpcw/icecube-submit) (version 111 is best LB.)",
    "2238320": "sinpcw thank you for the topic. Very interesting reading. And very good that you write about failure points. It is good information for beginners.",
    "2239181": "Thanks for sharing this topic. This will be helpful for beginners who are just starting out."
  },
  "source": "meta"
}