{
  "id": 266480,
  "title": "15th Place Solution : 8 days of Hard Work",
  "url": "/competitions/seti-breakthrough-listen/discussion/266480",
  "author_name": "Mr_KnowNothing",
  "post_date": "2021-08-19T09:17:16.886000",
  "votes": 38,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi all , <br>\nFirst of all I would like to thank Kaggle and organizers for their hard work to make this competition happen , the data after the reset was really good. We weren't sure of doing this competition after the reset as we were involved in other competitions but I would like to thank <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> and <a href=\"https://www.kaggle.com/shivamcyborg\" target=\"_blank\">@shivamcyborg</a> who joined me here , despite being tired after their SIIM Covid finish with<b> just 8 days left </b> </p>\n<h1>Summary</h1>\n<p>Our approach is pretty similar to others in the top solution , we just missed a few key points due to the short timeframe .</p>\n<p>Our modelling contains three phases :<br>\n1) Pretrain on Old Train Data <br>\n2) FineTune on New Train Data<br>\n3) Further tune on old Test Data and Pseudo labelled new Test Data</p>\n<p>We trained all models on the \"ON\" phase of data i.e [0,2,4] for quick iteration as the data was huge and time was less, also it seemed to work better for us</p>\n<h1>Phase 1 : Training</h1>\n<p>Before the competition reset , we were at 6th place without using the leak and had a lot of good strong models ,  So the first Phase of our training was already done for most models . We had logged everything properly which helped us to restart the competition quickly</p>\n<h1>Phase 2 Training</h1>\n<p>After Rejoining the competition we didn't train any model from scratch , we loaded all our models trained on Previous leaky Train Data and started training on the new data . </p>\n<p>We trained the following models in this stage and these all were used in our final ensemble:</p>\n<ul>\n<li>Effnet B5</li>\n<li>Effnet B7</li>\n<li>Eca-nfnet_l0</li>\n<li>Effnetv2m</li>\n</ul>\n<p>We used very light augs (vflip, hflip, randombrightnesscontrast, cutout) as we were afraid of losing the weak signals if we used heavy augs . All the models were trained for 20 epochs and used stratified five folds for CV</p>\n<h1>Phase 3 Training</h1>\n<p>We knew that the new train and test data had different distributions , so we started with pseudo on just highly confident labels and took only 5000 examples , this helped us increase CV but didn't reflect on LB , ofc we didn't give time to analyze why this was happening . We thought the old test data also might have some distribution similar to new test so finally we decided to retrain the models on all old test data and all new test data (pseudo labelled with our current best blend)</p>\n<p>In the third phase we trained the models with similar augs but with low LR and lower epoch</p>\n<h1>Quest Towards Finding the Magic</h1>\n<p>We were able to reach CV : 0.901 and LB : 0.795 with the aforementioned techniques and last few days were spent on finding the magic. From the hosts comment we all knew that there is a different set of signals in train , we used our best models embeddings to cluster and generate TSNE plot in order to isolate those images </p>\n<p><a href=\"https://ibb.co/WWnCy1K\"><img src=\"https://i.ibb.co/tMBRs5C/image-1.png\" alt=\"proc\"></a></p>\n<p><a href=\"https://ibb.co/ZNwxMg6\"><img src=\"https://i.ibb.co/SJ1B3mX/image-2.png\" alt=\"proc\"></a></p>\n<p>From the above plots it was clear to us that more than 50 percent of the test had little to no overlap with the train set , also we were able to see cluster 1 belongs mostly to test set and hence we started to look at the images in that and finally isolated the new signals .</p>\n<p><b> However We were not able to effectively use it </b></p>\n<h1>Things that didn't work us</h1>\n<ul>\n<li>Qishen ha's Alaska competition trick (Removing last few blocks of efficientnet and adding conv layers on top ) similar to <a href=\"https://www.kaggle.com/tanulsingh077/qishen-ha-alaska-implementation\" target=\"_blank\">this</a></li>\n<li>Reducing Model strides in conv stem to allow for longer training at high resolutions</li>\n<li>Residual Bi-LSTM heads</li>\n<li>Winning solution of previous SETI competition from <a href=\"https://github.com/sgrvinod/Wide-Residual-Nets-for-SETI?source=post_page\" target=\"_blank\">here</a></li>\n</ul>\n<h1>Lessons Learned</h1>\n<p>After reading up the top solutions , we missed up some really basic stuff probably due to lack of time and also due to some laziness . </p>\n<ul>\n<li>Mixup using the OR logic really makes sense</li>\n<li>Carefully analyzing the predictions/oofs (Error Analysis is one of the most important part of any machine learning problem)</li>\n</ul>\n<h1>Hardware used</h1>\n<p>The data in this competition was fairly huge , we used the following :</p>\n<ul>\n<li>Colab Pro (Mine)</li>\n<li>2xV100 ( <a href=\"https://www.kaggle.com/piantic\" target=\"_blank\">@piantic</a> )</li>\n<li>2xV100 (@shivamcyborg )</li>\n<li>2xA100 ( Rented from <a href=\"https://cloud.jarvislabs.ai/\" target=\"_blank\">https://cloud.jarvislabs.ai/</a>) </li>\n</ul>\n<p>I used jarvislabs cloud services for the first time and I found it very useful , especially the feature with which we can pause the instance and change the GPU type / cluster type without losing the hard disk or cpu.</p>\n<p>Similar to any other competition I am taking a good amount of learning from here as well , I really thank <a href=\"https://www.kaggle.com/markpeng\" target=\"_blank\">@markpeng</a> for his efforts in finding the magic , its always a pleasure working with him , I learned a lot , I also want to thank <a href=\"https://www.kaggle.com/piantic\" target=\"_blank\">@piantic</a> for all his efforts</p>",
  "messages": [
    {
      "id": 1481020,
      "postDate": "2021-08-19T09:17:16.887Z",
      "content": "<p>Hi all , <br>\nFirst of all I would like to thank Kaggle and organizers for their hard work to make this competition happen , the data after the reset was really good. We weren't sure of doing this competition after the reset as we were involved in other competitions but I would like to thank <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> and <a href=\"https://www.kaggle.com/shivamcyborg\" target=\"_blank\">@shivamcyborg</a> who joined me here , despite being tired after their SIIM Covid finish with<b> just 8 days left </b> </p>\n<h1>Summary</h1>\n<p>Our approach is pretty similar to others in the top solution , we just missed a few key points due to the short timeframe .</p>\n<p>Our modelling contains three phases :<br>\n1) Pretrain on Old Train Data <br>\n2) FineTune on New Train Data<br>\n3) Further tune on old Test Data and Pseudo labelled new Test Data</p>\n<p>We trained all models on the \"ON\" phase of data i.e [0,2,4] for quick iteration as the data was huge and time was less, also it seemed to work better for us</p>\n<h1>Phase 1 : Training</h1>\n<p>Before the competition reset , we were at 6th place without using the leak and had a lot of good strong models ,  So the first Phase of our training was already done for most models . We had logged everything properly which helped us to restart the competition quickly</p>\n<h1>Phase 2 Training</h1>\n<p>After Rejoining the competition we didn't train any model from scratch , we loaded all our models trained on Previous leaky Train Data and started training on the new data . </p>\n<p>We trained the following models in this stage and these all were used in our final ensemble:</p>\n<ul>\n<li>Effnet B5</li>\n<li>Effnet B7</li>\n<li>Eca-nfnet_l0</li>\n<li>Effnetv2m</li>\n</ul>\n<p>We used very light augs (vflip, hflip, randombrightnesscontrast, cutout) as we were afraid of losing the weak signals if we used heavy augs . All the models were trained for 20 epochs and used stratified five folds for CV</p>\n<h1>Phase 3 Training</h1>\n<p>We knew that the new train and test data had different distributions , so we started with pseudo on just highly confident labels and took only 5000 examples , this helped us increase CV but didn't reflect on LB , ofc we didn't give time to analyze why this was happening . We thought the old test data also might have some distribution similar to new test so finally we decided to retrain the models on all old test data and all new test data (pseudo labelled with our current best blend)</p>\n<p>In the third phase we trained the models with similar augs but with low LR and lower epoch</p>\n<h1>Quest Towards Finding the Magic</h1>\n<p>We were able to reach CV : 0.901 and LB : 0.795 with the aforementioned techniques and last few days were spent on finding the magic. From the hosts comment we all knew that there is a different set of signals in train , we used our best models embeddings to cluster and generate TSNE plot in order to isolate those images </p>\n<p><a href=\"https://ibb.co/WWnCy1K\"><img src=\"https://i.ibb.co/tMBRs5C/image-1.png\" alt=\"proc\"></a></p>\n<p><a href=\"https://ibb.co/ZNwxMg6\"><img src=\"https://i.ibb.co/SJ1B3mX/image-2.png\" alt=\"proc\"></a></p>\n<p>From the above plots it was clear to us that more than 50 percent of the test had little to no overlap with the train set , also we were able to see cluster 1 belongs mostly to test set and hence we started to look at the images in that and finally isolated the new signals .</p>\n<p><b> However We were not able to effectively use it </b></p>\n<h1>Things that didn't work us</h1>\n<ul>\n<li>Qishen ha's Alaska competition trick (Removing last few blocks of efficientnet and adding conv layers on top ) similar to <a href=\"https://www.kaggle.com/tanulsingh077/qishen-ha-alaska-implementation\" target=\"_blank\">this</a></li>\n<li>Reducing Model strides in conv stem to allow for longer training at high resolutions</li>\n<li>Residual Bi-LSTM heads</li>\n<li>Winning solution of previous SETI competition from <a href=\"https://github.com/sgrvinod/Wide-Residual-Nets-for-SETI?source=post_page\" target=\"_blank\">here</a></li>\n</ul>\n<h1>Lessons Learned</h1>\n<p>After reading up the top solutions , we missed up some really basic stuff probably due to lack of time and also due to some laziness . </p>\n<ul>\n<li>Mixup using the OR logic really makes sense</li>\n<li>Carefully analyzing the predictions/oofs (Error Analysis is one of the most important part of any machine learning problem)</li>\n</ul>\n<h1>Hardware used</h1>\n<p>The data in this competition was fairly huge , we used the following :</p>\n<ul>\n<li>Colab Pro (Mine)</li>\n<li>2xV100 ( <a href=\"https://www.kaggle.com/piantic\" target=\"_blank\">@piantic</a> )</li>\n<li>2xV100 (@shivamcyborg )</li>\n<li>2xA100 ( Rented from <a href=\"https://cloud.jarvislabs.ai/\" target=\"_blank\">https://cloud.jarvislabs.ai/</a>) </li>\n</ul>\n<p>I used jarvislabs cloud services for the first time and I found it very useful , especially the feature with which we can pause the instance and change the GPU type / cluster type without losing the hard disk or cpu.</p>\n<p>Similar to any other competition I am taking a good amount of learning from here as well , I really thank <a href=\"https://www.kaggle.com/markpeng\" target=\"_blank\">@markpeng</a> for his efforts in finding the magic , its always a pleasure working with him , I learned a lot , I also want to thank <a href=\"https://www.kaggle.com/piantic\" target=\"_blank\">@piantic</a> for all his efforts</p>",
      "rawMarkdown": "Hi all , \nFirst of all I would like to thank Kaggle and organizers for their hard work to make this competition happen , the data after the reset was really good. We weren't sure of doing this competition after the reset as we were involved in other competitions but I would like to thank @nischaydnk and @shivamcyborg who joined me here , despite being tired after their SIIM Covid finish with<b> just 8 days left </b> \n\n# Summary\n\nOur approach is pretty similar to others in the top solution , we just missed a few key points due to the short timeframe .\n\nOur modelling contains three phases :\n1) Pretrain on Old Train Data \n2) FineTune on New Train Data\n3) Further tune on old Test Data and Pseudo labelled new Test Data\n\nWe trained all models on the \"ON\" phase of data i.e [0,2,4] for quick iteration as the data was huge and time was less, also it seemed to work better for us\n\n# Phase 1 : Training \n\nBefore the competition reset , we were at 6th place without using the leak and had a lot of good strong models ,  So the first Phase of our training was already done for most models . We had logged everything properly which helped us to restart the competition quickly\n\n# Phase 2 Training\n\nAfter Rejoining the competition we didn't train any model from scratch , we loaded all our models trained on Previous leaky Train Data and started training on the new data . \n\nWe trained the following models in this stage and these all were used in our final ensemble:\n* Effnet B5\n* Effnet B7\n* Eca-nfnet_l0\n* Effnetv2m\n\nWe used very light augs (vflip, hflip, randombrightnesscontrast, cutout) as we were afraid of losing the weak signals if we used heavy augs . All the models were trained for 20 epochs and used stratified five folds for CV\n\n# Phase 3 Training\n\nWe knew that the new train and test data had different distributions , so we started with pseudo on just highly confident labels and took only 5000 examples , this helped us increase CV but didn't reflect on LB , ofc we didn't give time to analyze why this was happening . We thought the old test data also might have some distribution similar to new test so finally we decided to retrain the models on all old test data and all new test data (pseudo labelled with our current best blend)\n\nIn the third phase we trained the models with similar augs but with low LR and lower epoch\n\n# Quest Towards Finding the Magic\n\nWe were able to reach CV : 0.901 and LB : 0.795 with the aforementioned techniques and last few days were spent on finding the magic. From the hosts comment we all knew that there is a different set of signals in train , we used our best models embeddings to cluster and generate TSNE plot in order to isolate those images \n\n<a href=\"https://ibb.co/WWnCy1K\"><img src=\"https://i.ibb.co/tMBRs5C/image-1.png\" alt=\"proc\" border=\"0\"></a>\n\n\n<a href=\"https://ibb.co/ZNwxMg6\"><img src=\"https://i.ibb.co/SJ1B3mX/image-2.png\" alt=\"proc\" border=\"0\"></a>\n\nFrom the above plots it was clear to us that more than 50 percent of the test had little to no overlap with the train set , also we were able to see cluster 1 belongs mostly to test set and hence we started to look at the images in that and finally isolated the new signals .\n\n<b> However We were not able to effectively use it </b>\n\n# Things that didn't work us\n\n* Qishen ha's Alaska competition trick (Removing last few blocks of efficientnet and adding conv layers on top ) similar to [this](https://www.kaggle.com/tanulsingh077/qishen-ha-alaska-implementation)\n* Reducing Model strides in conv stem to allow for longer training at high resolutions\n* Residual Bi-LSTM heads\n* Winning solution of previous SETI competition from [here](https://github.com/sgrvinod/Wide-Residual-Nets-for-SETI?source=post_page)\n\n# Lessons Learned\n\nAfter reading up the top solutions , we missed up some really basic stuff probably due to lack of time and also due to some laziness . \n* Mixup using the OR logic really makes sense\n* Carefully analyzing the predictions/oofs (Error Analysis is one of the most important part of any machine learning problem)\n\n# Hardware used \nThe data in this competition was fairly huge , we used the following :\n* Colab Pro (Mine)\n* 2xV100 ( @piantic )\n* 2xV100 (@shivamcyborg )\n* 2xA100 ( Rented from https://cloud.jarvislabs.ai/) \n\nI used jarvislabs cloud services for the first time and I found it very useful , especially the feature with which we can pause the instance and change the GPU type / cluster type without losing the hard disk or cpu.\n\nSimilar to any other competition I am taking a good amount of learning from here as well , I really thank @markpeng for his efforts in finding the magic , its always a pleasure working with him , I learned a lot , I also want to thank @piantic for all his efforts",
      "votes": 38
    },
    {
      "id": 1495851,
      "postDate": "2021-08-29T20:43:13.243Z",
      "content": "<p>Good job =))</p>",
      "rawMarkdown": "Good job =))\n",
      "votes": 1
    },
    {
      "id": 1481043,
      "postDate": "2021-08-19T09:33:19.507Z",
      "content": "<p>Great work! I learned a lot from you.<br>\nYour method is clear and straightforward to understand.</p>",
      "rawMarkdown": "Great work! I learned a lot from you.\nYour method is clear and straightforward to understand.",
      "votes": 2,
      "replies": [
        {
          "id": 1481886,
          "postDate": "2021-08-19T17:52:55.367Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/woosungyoon\" target=\"_blank\">@woosungyoon</a> </p>",
          "rawMarkdown": "Thanks @woosungyoon "
        }
      ]
    },
    {
      "id": 1500773,
      "postDate": "2021-09-02T16:06:47.850Z",
      "content": "<p>You are really inspirational for Kagglers like me!!</p>\n<p>Excellent work <a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> Keep it up bro.</p>",
      "rawMarkdown": "You are really inspirational for Kagglers like me!!\n\nExcellent work @tanulsingh077 Keep it up bro."
    }
  ],
  "comments": [
    {
      "id": 1495851,
      "author_name": "Fuco",
      "author_url": "",
      "post_date": "2021-08-29T20:43:13.243000",
      "content": "<p>Good job =))</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1481043,
      "author_name": "WOOSUNG YOON",
      "author_url": "",
      "post_date": "2021-08-19T09:33:19.507000",
      "content": "<p>Great work! I learned a lot from you.<br>\nYour method is clear and straightforward to understand.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1481886,
          "author_name": "Mr_KnowNothing",
          "author_url": "",
          "post_date": "2021-08-19T17:52:55.367000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/woosungyoon\" target=\"_blank\">@woosungyoon</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1500773,
      "author_name": "Vinayak Shanawad",
      "author_url": "",
      "post_date": "2021-09-02T16:06:47.850000",
      "content": "<p>You are really inspirational for Kagglers like me!!</p>\n<p>Excellent work <a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> Keep it up bro.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1481020": "Hi all , \nFirst of all I would like to thank Kaggle and organizers for their hard work to make this competition happen , the data after the reset was really good. We weren't sure of doing this competition after the reset as we were involved in other competitions but I would like to thank @nischaydnk and @shivamcyborg who joined me here , despite being tired after their SIIM Covid finish with<b> just 8 days left </b> \n\n# Summary\n\nOur approach is pretty similar to others in the top solution , we just missed a few key points due to the short timeframe .\n\nOur modelling contains three phases :\n1) Pretrain on Old Train Data \n2) FineTune on New Train Data\n3) Further tune on old Test Data and Pseudo labelled new Test Data\n\nWe trained all models on the \"ON\" phase of data i.e [0,2,4] for quick iteration as the data was huge and time was less, also it seemed to work better for us\n\n# Phase 1 : Training \n\nBefore the competition reset , we were at 6th place without using the leak and had a lot of good strong models ,  So the first Phase of our training was already done for most models . We had logged everything properly which helped us to restart the competition quickly\n\n# Phase 2 Training\n\nAfter Rejoining the competition we didn't train any model from scratch , we loaded all our models trained on Previous leaky Train Data and started training on the new data . \n\nWe trained the following models in this stage and these all were used in our final ensemble:\n* Effnet B5\n* Effnet B7\n* Eca-nfnet_l0\n* Effnetv2m\n\nWe used very light augs (vflip, hflip, randombrightnesscontrast, cutout) as we were afraid of losing the weak signals if we used heavy augs . All the models were trained for 20 epochs and used stratified five folds for CV\n\n# Phase 3 Training\n\nWe knew that the new train and test data had different distributions , so we started with pseudo on just highly confident labels and took only 5000 examples , this helped us increase CV but didn't reflect on LB , ofc we didn't give time to analyze why this was happening . We thought the old test data also might have some distribution similar to new test so finally we decided to retrain the models on all old test data and all new test data (pseudo labelled with our current best blend)\n\nIn the third phase we trained the models with similar augs but with low LR and lower epoch\n\n# Quest Towards Finding the Magic\n\nWe were able to reach CV : 0.901 and LB : 0.795 with the aforementioned techniques and last few days were spent on finding the magic. From the hosts comment we all knew that there is a different set of signals in train , we used our best models embeddings to cluster and generate TSNE plot in order to isolate those images \n\n<a href=\"https://ibb.co/WWnCy1K\"><img src=\"https://i.ibb.co/tMBRs5C/image-1.png\" alt=\"proc\" border=\"0\"></a>\n\n\n<a href=\"https://ibb.co/ZNwxMg6\"><img src=\"https://i.ibb.co/SJ1B3mX/image-2.png\" alt=\"proc\" border=\"0\"></a>\n\nFrom the above plots it was clear to us that more than 50 percent of the test had little to no overlap with the train set , also we were able to see cluster 1 belongs mostly to test set and hence we started to look at the images in that and finally isolated the new signals .\n\n<b> However We were not able to effectively use it </b>\n\n# Things that didn't work us\n\n* Qishen ha's Alaska competition trick (Removing last few blocks of efficientnet and adding conv layers on top ) similar to [this](https://www.kaggle.com/tanulsingh077/qishen-ha-alaska-implementation)\n* Reducing Model strides in conv stem to allow for longer training at high resolutions\n* Residual Bi-LSTM heads\n* Winning solution of previous SETI competition from [here](https://github.com/sgrvinod/Wide-Residual-Nets-for-SETI?source=post_page)\n\n# Lessons Learned\n\nAfter reading up the top solutions , we missed up some really basic stuff probably due to lack of time and also due to some laziness . \n* Mixup using the OR logic really makes sense\n* Carefully analyzing the predictions/oofs (Error Analysis is one of the most important part of any machine learning problem)\n\n# Hardware used \nThe data in this competition was fairly huge , we used the following :\n* Colab Pro (Mine)\n* 2xV100 ( @piantic )\n* 2xV100 (@shivamcyborg )\n* 2xA100 ( Rented from https://cloud.jarvislabs.ai/) \n\nI used jarvislabs cloud services for the first time and I found it very useful , especially the feature with which we can pause the instance and change the GPU type / cluster type without losing the hard disk or cpu.\n\nSimilar to any other competition I am taking a good amount of learning from here as well , I really thank @markpeng for his efforts in finding the magic , its always a pleasure working with him , I learned a lot , I also want to thank @piantic for all his efforts",
    "1495851": "Good job =))\n",
    "1481043": "Great work! I learned a lot from you.\nYour method is clear and straightforward to understand.",
    "1500773": "You are really inspirational for Kagglers like me!!\n\nExcellent work @tanulsingh077 Keep it up bro."
  }
}