{
  "id": 275334,
  "title": "5th place solution (instead of)",
  "url": "/competitions/g2net-gravitational-wave-detection/writeups/gleb-5th-place-solution-instead-of",
  "author_name": "",
  "post_date": "2021-09-30T01:19:49.377657300Z",
  "votes": 30,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Thanks to kaggle, organizers, and community for intense battle at the end and live discussion in the beginning.</p>\n<p>Not really a solution post, rather rant on EfficientNets. Solution itself is pretty straightforward, just 2d CNN go brrrrrr.</p>\n<h1>TLDR</h1>\n<ul>\n<li>PSD whitening with precomputed mean PSD for each channel, custom window function : mirrored sigmoid</li>\n<li>1d augs: shift, noise, mixup, cutmix, shuffle</li>\n<li>2d CQT default settings fmin=16, fmax=1024, hop=12</li>\n<li>5xCNN 2d ensemble with TTA, models and EMA's</li>\n<li>EfficientNets are not so good (for me).</li>\n</ul>\n<h1>Intro</h1>\n<p>As i don't know anything about competition problem and do not have any experience at 1d-audio-spectrogram-ish type of tasks: I decided to go full research mode on specific NN architecture</p>\n<p>Disclaimer: CNN shouldn't be the right fit for spectrograms, but they are. I also tried Transformers, surely they can work, but score-wise i was barely able to <br>\ngo into 88.+ on LB. I'm still pretty new with transformers thou.</p>\n<h1>EfficientNets</h1>\n<p>What a original choice, you might say. Indeed, a lot of people uses EN for some reason, but not me. I don't use EN family because it is not … efficient.<br>\nIt is not much of a secret and there are some papers that are pointing at that (i.e. RegNet, GENet, EffnetV2, etc). But what if I was wrong all that time and <br>\nimg/sec -to- acc ratio is much better? </p>\n<p>It is not.</p>\n<p>My best result with EN after a lot of tuning and fitting was easily beaten (not by much) on the first try with RegNet and even ResNet. <strong>And they are 4x times faster.</strong><br>\nOf course on some datasets EN will work better, but im still to find one, other then ImageNet1k.<br>\nThere can be a lot of reasons why it is what is is, my thoughts on that: it is happening because EN is \"overfitting\" ImageNet1k (can be NAS or compound-scaling)</p>\n<p>My best EN ensemble score was 8830 LB, after that I gave up and get ResNets into the mix in last couple of days of the competition.</p>\n<p>Thanks for reading and please consider some other arch's if you dont already.</p>",
  "messages": [
    {
      "id": "1528813",
      "postDate": "09/30/2021 01:19:49",
      "content": "<p>Thanks to kaggle, organizers, and community for intense battle at the end and live discussion in the beginning.</p>\n<p>Not really a solution post, rather rant on EfficientNets. Solution itself is pretty straightforward, just 2d CNN go brrrrrr.</p>\n<h1>TLDR</h1>\n<ul>\n<li>PSD whitening with precomputed mean PSD for each channel, custom window function : mirrored sigmoid</li>\n<li>1d augs: shift, noise, mixup, cutmix, shuffle</li>\n<li>2d CQT default settings fmin=16, fmax=1024, hop=12</li>\n<li>5xCNN 2d ensemble with TTA, models and EMA's</li>\n<li>EfficientNets are not so good (for me).</li>\n</ul>\n<h1>Intro</h1>\n<p>As i don't know anything about competition problem and do not have any experience at 1d-audio-spectrogram-ish type of tasks: I decided to go full research mode on specific NN architecture</p>\n<p>Disclaimer: CNN shouldn't be the right fit for spectrograms, but they are. I also tried Transformers, surely they can work, but score-wise i was barely able to <br>\ngo into 88.+ on LB. I'm still pretty new with transformers thou.</p>\n<h1>EfficientNets</h1>\n<p>What a original choice, you might say. Indeed, a lot of people uses EN for some reason, but not me. I don't use EN family because it is not … efficient.<br>\nIt is not much of a secret and there are some papers that are pointing at that (i.e. RegNet, GENet, EffnetV2, etc). But what if I was wrong all that time and <br>\nimg/sec -to- acc ratio is much better? </p>\n<p>It is not.</p>\n<p>My best result with EN after a lot of tuning and fitting was easily beaten (not by much) on the first try with RegNet and even ResNet. <strong>And they are 4x times faster.</strong><br>\nOf course on some datasets EN will work better, but im still to find one, other then ImageNet1k.<br>\nThere can be a lot of reasons why it is what is is, my thoughts on that: it is happening because EN is \"overfitting\" ImageNet1k (can be NAS or compound-scaling)</p>\n<p>My best EN ensemble score was 8830 LB, after that I gave up and get ResNets into the mix in last couple of days of the competition.</p>\n<p>Thanks for reading and please consider some other arch's if you dont already.</p>",
      "rawMarkdown": "Thanks to kaggle, organizers, and community for intense battle at the end and live discussion in the beginning.\n\nNot really a solution post, rather rant on EfficientNets. Solution itself is pretty straightforward, just 2d CNN go brrrrrr.\n\n# TLDR\n\n- PSD whitening with precomputed mean PSD for each channel, custom window function : mirrored sigmoid\n- 1d augs: shift, noise, mixup, cutmix, shuffle\n- 2d CQT default settings fmin=16, fmax=1024, hop=12\n- 5xCNN 2d ensemble with TTA, models and EMA's\n- EfficientNets are not so good (for me).\n\n# Intro\n\nAs i don't know anything about competition problem and do not have any experience at 1d-audio-spectrogram-ish type of tasks: I decided to go full research mode on specific NN architecture\n\nDisclaimer: CNN shouldn't be the right fit for spectrograms, but they are. I also tried Transformers, surely they can work, but score-wise i was barely able to \ngo into 88.+ on LB. I'm still pretty new with transformers thou.\n\n# EfficientNets\n\nWhat a original choice, you might say. Indeed, a lot of people uses EN for some reason, but not me. I don't use EN family because it is not ... efficient.\nIt is not much of a secret and there are some papers that are pointing at that (i.e. RegNet, GENet, EffnetV2, etc). But what if I was wrong all that time and \nimg/sec -to- acc ratio is much better? \n\nIt is not.\n\nMy best result with EN after a lot of tuning and fitting was easily beaten (not by much) on the first try with RegNet and even ResNet. **And they are 4x times faster.**\nOf course on some datasets EN will work better, but im still to find one, other then ImageNet1k.\nThere can be a lot of reasons why it is what is is, my thoughts on that: it is happening because EN is \"overfitting\" ImageNet1k (can be NAS or compound-scaling)\n\nMy best EN ensemble score was 8830 LB, after that I gave up and get ResNets into the mix in last couple of days of the competition.\n\nThanks for reading and please consider some other arch's if you dont already.",
      "votes": null
    },
    {
      "id": "1529069",
      "postDate": "09/30/2021 06:14:19",
      "content": "<p>resnest26d was also good architecture.. just that it was not so memory efficient and  your preprocessing worked well for you. i think that is more dersirable here ,most of works on gw focus on Preprocessing rather modelling .Could you throw more light on how u computed the PSD . Like you computed it for whole dataset and used ,or  signal wise PSD. I tried but it dint work</p>",
      "rawMarkdown": "resnest26d was also good architecture.. just that it was not so memory efficient and  your preprocessing worked well for you. i think that is more dersirable here ,most of works on gw focus on Preprocessing rather modelling .Could you throw more light on how u computed the PSD . Like you computed it for whole dataset and used ,or  signal wise PSD. I tried but it dint work",
      "votes": null
    },
    {
      "id": "1529198",
      "postDate": "09/30/2021 08:12:41",
      "content": "<p>Congrats.  I don't know why but I could not get resnet34 work better than effnet v2s.  Maybe we are the only team failing to use resnets here.</p>",
      "rawMarkdown": "Congrats.  I don't know why but I could not get resnet34 work better than effnet v2s.  Maybe we are the only team failing to use resnets here.",
      "votes": null
    },
    {
      "id": "1559880",
      "postDate": "10/27/2021 08:05:24",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1529069,
      "author_name": "jaideepvalani",
      "author_url": "",
      "post_date": "09/30/2021 06:14:19",
      "content": "<p>resnest26d was also good architecture.. just that it was not so memory efficient and  your preprocessing worked well for you. i think that is more dersirable here ,most of works on gw focus on Preprocessing rather modelling .Could you throw more light on how u computed the PSD . Like you computed it for whole dataset and used ,or  signal wise PSD. I tried but it dint work</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1529198,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "09/30/2021 08:12:41",
      "content": "<p>Congrats.  I don't know why but I could not get resnet34 work better than effnet v2s.  Maybe we are the only team failing to use resnets here.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1559880,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 08:05:24",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1528813": "Thanks to kaggle, organizers, and community for intense battle at the end and live discussion in the beginning.\n\nNot really a solution post, rather rant on EfficientNets. Solution itself is pretty straightforward, just 2d CNN go brrrrrr.\n\n# TLDR\n\n- PSD whitening with precomputed mean PSD for each channel, custom window function : mirrored sigmoid\n- 1d augs: shift, noise, mixup, cutmix, shuffle\n- 2d CQT default settings fmin=16, fmax=1024, hop=12\n- 5xCNN 2d ensemble with TTA, models and EMA's\n- EfficientNets are not so good (for me).\n\n# Intro\n\nAs i don't know anything about competition problem and do not have any experience at 1d-audio-spectrogram-ish type of tasks: I decided to go full research mode on specific NN architecture\n\nDisclaimer: CNN shouldn't be the right fit for spectrograms, but they are. I also tried Transformers, surely they can work, but score-wise i was barely able to \ngo into 88.+ on LB. I'm still pretty new with transformers thou.\n\n# EfficientNets\n\nWhat a original choice, you might say. Indeed, a lot of people uses EN for some reason, but not me. I don't use EN family because it is not ... efficient.\nIt is not much of a secret and there are some papers that are pointing at that (i.e. RegNet, GENet, EffnetV2, etc). But what if I was wrong all that time and \nimg/sec -to- acc ratio is much better? \n\nIt is not.\n\nMy best result with EN after a lot of tuning and fitting was easily beaten (not by much) on the first try with RegNet and even ResNet. **And they are 4x times faster.**\nOf course on some datasets EN will work better, but im still to find one, other then ImageNet1k.\nThere can be a lot of reasons why it is what is is, my thoughts on that: it is happening because EN is \"overfitting\" ImageNet1k (can be NAS or compound-scaling)\n\nMy best EN ensemble score was 8830 LB, after that I gave up and get ResNets into the mix in last couple of days of the competition.\n\nThanks for reading and please consider some other arch's if you dont already.",
    "1529069": "resnest26d was also good architecture.. just that it was not so memory efficient and  your preprocessing worked well for you. i think that is more dersirable here ,most of works on gw focus on Preprocessing rather modelling .Could you throw more light on how u computed the PSD . Like you computed it for whole dataset and used ,or  signal wise PSD. I tried but it dint work",
    "1529198": "Congrats.  I don't know why but I could not get resnet34 work better than effnet v2s.  Maybe we are the only team failing to use resnets here.",
    "1559880": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris"
  },
  "source": "meta"
}