{
  "id": 271576,
  "title": "class activation map again ... for 1d CNN model",
  "url": "/competitions/g2net-gravitational-wave-detection/discussion/271576",
  "author_name": "",
  "post_date": "2021-09-11T09:10:14.035616Z",
  "votes": 39,
  "comment_count": 32,
  "views": 0,
  "content": "<p>this is a preview only.<br>\nmore analysis and results coming up soon!</p>\n<p><img src=\"https://i.ibb.co/7Cm6Fwy/Selection-839.png\" alt=\"https://i.ibb.co/7Cm6Fwy/Selection-839.png\"></p>",
  "messages": [
    {
      "id": "1509426",
      "postDate": "09/11/2021 09:10:14",
      "content": "<p>this is a preview only.<br>\nmore analysis and results coming up soon!</p>\n<p><img src=\"https://i.ibb.co/7Cm6Fwy/Selection-839.png\" alt=\"https://i.ibb.co/7Cm6Fwy/Selection-839.png\"></p>",
      "rawMarkdown": "this is a preview only.\nmore analysis and results coming up soon!\n\n![https://i.ibb.co/7Cm6Fwy/Selection-839.png](https://i.ibb.co/7Cm6Fwy/Selection-839.png)",
      "votes": null
    },
    {
      "id": "1509905",
      "postDate": "09/11/2021 19:13:51",
      "content": "<blockquote>\n  <p>The CQT is not used in the CNN</p>\n</blockquote>\n<p>What's used ?</p>",
      "rawMarkdown": "> The CQT is not used in the CNN\n\nWhat's used ?",
      "votes": null
    },
    {
      "id": "1509910",
      "postDate": "09/11/2021 19:17:07",
      "content": "<p>It's a 1d cnn, you're thinking of 2d cnn. In this case, he probably used the raw time series.</p>",
      "rawMarkdown": "It's a 1d cnn, you're thinking of 2d cnn. In this case, he probably used the raw time series.",
      "votes": null
    },
    {
      "id": "1509912",
      "postDate": "09/11/2021 19:19:16",
      "content": "<p>Oops sorry, missed that part !</p>",
      "rawMarkdown": "Oops sorry, missed that part !",
      "votes": null
    },
    {
      "id": "1510028",
      "postDate": "09/12/2021 01:09:36",
      "content": "<p>update1.<br>\n(selection of max kernel (red box) could be wrong as I did not consider the weights of the next bn layer)</p>\n<p><img src=\"https://i.ibb.co/n3gQzy8/Selection-851.png\" alt=\"https://i.ibb.co/n3gQzy8/Selection-851.png\"><br>\n<img src=\"https://i.ibb.co/C70jXPQ/Selection-850.png\" alt=\"https://i.ibb.co/C70jXPQ/Selection-850.png\"></p>",
      "rawMarkdown": "update1.\n(selection of max kernel (red box) could be wrong as I did not consider the weights of the next bn layer)\n\n![https://i.ibb.co/n3gQzy8/Selection-851.png](https://i.ibb.co/n3gQzy8/Selection-851.png)\n![https://i.ibb.co/C70jXPQ/Selection-850.png](https://i.ibb.co/C70jXPQ/Selection-850.png)",
      "votes": null
    },
    {
      "id": "1510029",
      "postDate": "09/12/2021 01:17:44",
      "content": "<p>if you are interested in how to design waveform 1d CNN and compare it with spectrogram 2d CNN, refer to:</p>\n<p>ISMIR 2019 tutorial: waveform-based music processing with deep learning<br>\n<a href=\"https://zenodo.org/record/3529714#.YT1UZnUzbCJ\" target=\"_blank\">https://zenodo.org/record/3529714#.YT1UZnUzbCJ</a></p>\n<p>Learning Multiscale Features Directly From Waveforms<br>\n<a href=\"http://research.baidu.com/Public/uploads/5ac04fa6a3f28.PDF\" target=\"_blank\">http://research.baidu.com/Public/uploads/5ac04fa6a3f28.PDF</a></p>\n<p>SampleCNN: End-to-End Deep Convolutional Neural Networks Using Very Small Filters for Music Classification</p>",
      "rawMarkdown": "if you are interested in how to design waveform 1d CNN and compare it with spectrogram 2d CNN, refer to:\n\nISMIR 2019 tutorial: waveform-based music processing with deep learning\nhttps://zenodo.org/record/3529714#.YT1UZnUzbCJ\n\nLearning Multiscale Features Directly From Waveforms\nhttp://research.baidu.com/Public/uploads/5ac04fa6a3f28.PDF\n\n\nSampleCNN: End-to-End Deep Convolutional Neural Networks Using Very Small Filters for Music Classification",
      "votes": null
    },
    {
      "id": "1510036",
      "postDate": "09/12/2021 02:11:13",
      "content": "<p>Interesting. It seems as though you're running your 1d cnn on frequency domain magnitudes as opposed to time domain?</p>",
      "rawMarkdown": "Interesting. It seems as though you're running your 1d cnn on frequency domain magnitudes as opposed to time domain?",
      "votes": null
    },
    {
      "id": "1510038",
      "postDate": "09/12/2021 02:16:53",
      "content": "<p>no. i am running on time domain</p>",
      "rawMarkdown": "no. i am running on time domain",
      "votes": null
    },
    {
      "id": "1511136",
      "postDate": "09/13/2021 06:43:39",
      "content": "<p>interesting, whitening don't work for both 1d cnn and 2d cnn.<br>\n(maybe 2 sec is too short to estimate noise power)</p>\n<p>i transfer the 2d-cnn signal processing for 1d-cnn, the performance of 1d-cnn is greatly increased</p>\n<p>simple 8 layer 1dcnn has cv = 0.868902<br>\n(45 min to train up to 6 epoach)</p>\n<p>the same wave signal when transform with CQT(128x256)+eff-b0 gives cv = 0.8716</p>\n<p>ensembling gives  cv =0.873843</p>\n<hr>\n<p>bandpass seems to be a must for 1d-cnn</p>",
      "rawMarkdown": "interesting, whitening don't work for both 1d cnn and 2d cnn.\n(maybe 2 sec is too short to estimate noise power)\n\ni transfer the 2d-cnn signal processing for 1d-cnn, the performance of 1d-cnn is greatly increased\n\nsimple 8 layer 1dcnn has cv = 0.868902\n(45 min to train up to 6 epoach)\n\nthe same wave signal when transform with CQT(128x256)+eff-b0 gives cv = 0.8716\n\nensembling gives  cv =0.873843\n\n---\n\nbandpass seems to be a must for 1d-cnn",
      "votes": null
    },
    {
      "id": "1511166",
      "postDate": "09/13/2021 07:15:39",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Are you saying you are feeding raw waveform + bandpass to 1D CNNs? Any other pre-processing that you can tell us about?  </p>",
      "rawMarkdown": "hengck23 Are you saying you are feeding raw waveform + bandpass to 1D CNNs? Any other pre-processing that you can tell us about?",
      "votes": null
    },
    {
      "id": "1511309",
      "postDate": "09/13/2021 10:11:16",
      "content": "<p>after increasing the parameters of the 1d-cnn:<br>\n1d-cnn auc  cv = 0.870425<br>\nensemble auc cv = 0.874456</p>\n<p></p>\n<p>(this is wrong. i forget that the y-axis is log scale)</p>\n<p><img src=\"https://i.ibb.co/xhFrMYs/Selection-864.png\" alt=\"https://i.ibb.co/xhFrMYs/Selection-864.png\"><br>\n<img src=\"https://i.ibb.co/fCX7YLw/Selection-863.png\" alt=\"https://i.ibb.co/fCX7YLw/Selection-863.png\"></p>\n<hr>\n<p>instead of ensembling, why not train jointly (or partially jointly, e.g. freeze one of the network if you have limited GPU memory)?</p>\n<p>if the 1d-CNN of the ensemble gets accurate enough, one might want to try distillation</p>",
      "rawMarkdown": "after increasing the parameters of the 1d-cnn:\n1d-cnn auc  cv = 0.870425\nensemble auc cv = 0.874456\n\n\n~~where did the improvement come from?\ni think we did not detect more GW. But we did reduce the number of FP~~\n\n(this is wrong. i forget that the y-axis is log scale)\n\n\n![https://i.ibb.co/xhFrMYs/Selection-864.png](https://i.ibb.co/xhFrMYs/Selection-864.png)\n![https://i.ibb.co/fCX7YLw/Selection-863.png](https://i.ibb.co/fCX7YLw/Selection-863.png)\n\n---\ninstead of ensembling, why not train jointly (or partially jointly, e.g. freeze one of the network if you have limited GPU memory)?\n\nif the 1d-CNN of the ensemble gets accurate enough, one might want to try distillation",
      "votes": null
    },
    {
      "id": "1511470",
      "postDate": "09/13/2021 12:58:52",
      "content": "<blockquote>\n  <p>instead of ensembling, why not train jointly</p>\n</blockquote>\n<p>I did this. But 1d cnn learning schedule is different from 2d, so either separate training then joining needs to be done. Or for more optimal results, careful finetuning of differential learning rates, which I haven't explored.</p>",
      "rawMarkdown": "> instead of ensembling, why not train jointly\n\nI did this. But 1d cnn learning schedule is different from 2d, so either separate training then joining needs to be done. Or for more optimal results, careful finetuning of differential learning rates, which I haven't explored.",
      "votes": null
    },
    {
      "id": "1511500",
      "postDate": "09/13/2021 13:27:41",
      "content": "<p>i just use trained 1d and 2d network as feature extractors.<br>\nthey are freeze when training a combined network.</p>\n<p>but even so, there seems to be overfitting, so maybe it is not easy to train jointly</p>\n<p>we can flip left-right and up-down for noise wave form. can we do that for GW samples as well?</p>",
      "rawMarkdown": "i just use trained 1d and 2d network as feature extractors.\nthey are freeze when training a combined network.\n\nbut even so, there seems to be overfitting, so maybe it is not easy to train jointly\n\n\nwe can flip left-right and up-down for noise wave form. can we do that for GW samples as well?",
      "votes": null
    },
    {
      "id": "1511550",
      "postDate": "09/13/2021 14:09:41",
      "content": "<p>I'm not sure that effect of reduction of FP comes from exactly \"ensemble of 1D-net and 2D-net\" but instead because of the nature of ensemble itself<br>\nWhat I'm saying is that you could achieve same result with ensembling using only 2D-nets, I suppose<br>\nIt would be interesting to look on the same plots, but with ensemble of 2d-nets only, I should try to do that myself I suppose</p>\n<p>By the way, how to you define \"pos\" or \"neg\" in your plots? <br>\nBecause usually we compute ROC-AUC on the raw probabilities, but judging from the second plot you divided samples into \"pos/neg\" somehow.</p>",
      "rawMarkdown": "I'm not sure that effect of reduction of FP comes from exactly \"ensemble of 1D-net and 2D-net\" but instead because of the nature of ensemble itself\nWhat I'm saying is that you could achieve same result with ensembling using only 2D-nets, I suppose\nIt would be interesting to look on the same plots, but with ensemble of 2d-nets only, I should try to do that myself I suppose\n\nBy the way, how to you define \"pos\" or \"neg\" in your plots? \nBecause usually we compute ROC-AUC on the raw probabilities, but judging from the second plot you divided samples into \"pos/neg\" somehow.",
      "votes": null
    },
    {
      "id": "1511551",
      "postDate": "09/13/2021 14:11:56",
      "content": "<p>By the way, how to you define \"pos\" or \"neg\" in your plots?<br>\n…</p>\n<p>we have the ground truth label in validation for the distribution plots.<br>\nfor ROC, it is for all train samples (i.e. pos and neg)</p>",
      "rawMarkdown": "By the way, how to you define \"pos\" or \"neg\" in your plots?\n...\n\nwe have the ground truth label in validation for the distribution plots.\nfor ROC, it is for all train samples (i.e. pos and neg)",
      "votes": null
    },
    {
      "id": "1511565",
      "postDate": "09/13/2021 14:19:09",
      "content": "<p><a href=\"https://www.kaggle.com/rimuru\" target=\"_blank\">@rimuru</a> try something like:</p>\n<pre><code>hist, edges = np.histogram(z.probs, bins=100)\nplt.plot(edges[:-1], hist, linewidth=1)\nplt.yscale('log')\nplt.show()\n</code></pre>\n<p>But split z by train target or oof target.</p>",
      "rawMarkdown": "rimuru try something like:\n\n```\nhist, edges = np.histogram(z.probs, bins=100)\nplt.plot(edges[:-1], hist, linewidth=1)\nplt.yscale('log')\nplt.show()\n```\n\nBut split z by train target or oof target.",
      "votes": null
    },
    {
      "id": "1511811",
      "postDate": "09/13/2021 17:52:39",
      "content": "<p>after more augmentations:</p>\n<p></p>\n<p>for one fold only:</p>\n<p>2d-cnn ( CQT 128x256 + eff-b0)  auc cv = 0.872293 <br>\n1d-cnn (wave input to 8 layer temporal conv) auc cv = 0.873005 (lb =0.8746)<br>\nensemble:  0.875676 (this is almost the same results as effnet b7)</p>\n<p>conclusion:</p>\n<ol>\n<li>it is possible that 1d-cnn outperforms 2d-cnn</li>\n<li>you can get good results using low computation model</li>\n<li>it is possible to train from scratch</li>\n</ol>",
      "rawMarkdown": "after more augmentations:\n\n~~1d-cnn auc cv = 0.872519\nensemble auc cv = 0.875179  (this is almost the same results as effnet b7)~~\n\nfor one fold only:\n\n2d-cnn ( CQT 128x256 + eff-b0)  auc cv = 0.872293 \n1d-cnn (wave input to 8 layer temporal conv) auc cv = 0.873005 (lb =0.8746)\nensemble:  0.875676 (this is almost the same results as effnet b7)\n\n\nconclusion:\n1. it is possible that 1d-cnn outperforms 2d-cnn\n2. you can get good results using low computation model\n3. it is possible to train from scratch",
      "votes": null
    },
    {
      "id": "1511850",
      "postDate": "09/13/2021 18:18:45",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> are you using the 1d-cnn on the signal directly or on the spectrogram ?</p>",
      "rawMarkdown": "hengck23 are you using the 1d-cnn on the signal directly or on the spectrogram ?",
      "votes": null
    },
    {
      "id": "1511860",
      "postDate": "09/13/2021 18:23:32",
      "content": "<pre><code>class Net(nn.Module):\n    # 1D convolutional neural network.   \n    def __init__(self):\n        super().__init__()\n        self.cnn1 =  nnn.Conv1d(3, 64, kernel_size=64, padding=32)\n       ... etc ...\n\n\n    def forward(self, wave):\n        x = self.cnn1(wave)  \n       ... etc ...\n\n\ndef run_check_net():\n    batch_size = 5\n    wave = torch.randn(batch_size, 3, 4096) \n\n    net = Net() \n    signal= net(wave)\n</code></pre>",
      "rawMarkdown": "```\nclass Net(nn.Module):\n    # 1D convolutional neural network.   \n    def __init__(self):\n        super().__init__()\n        self.cnn1 =  nnn.Conv1d(3, 64, kernel_size=64, padding=32)\n       ... etc ...\n     \n\n    def forward(self, wave):\n        x = self.cnn1(wave)  \n       ... etc ...\n\n\ndef run_check_net():\n    batch_size = 5\n    wave = torch.randn(batch_size, 3, 4096) \n\n    net = Net() \n    signal= net(wave)\n\n```",
      "votes": null
    },
    {
      "id": "1512153",
      "postDate": "09/14/2021 03:21:05",
      "content": "<p><img src=\"https://i.ibb.co/5F3dvjD/Selection-868.png\" alt=\"https://i.ibb.co/5F3dvjD/Selection-868.png\"></p>\n<p>my guess.</p>\n<p>do note that performance is limited by data.<br>\nanything that is in the \"front\" (e.g. data, augmentation, signal processing) set the limit of performance.<br>\nanything in the \"back\" determines how close you can reach the limit</p>",
      "rawMarkdown": "![https://i.ibb.co/5F3dvjD/Selection-868.png](https://i.ibb.co/5F3dvjD/Selection-868.png)\n\nmy guess.\n\ndo note that performance is limited by data.\nanything that is in the \"front\" (e.g. data, augmentation, signal processing) set the limit of performance.\nanything in the \"back\" determines how close you can reach the limit",
      "votes": null
    },
    {
      "id": "1512264",
      "postDate": "09/14/2021 06:20:48",
      "content": "<p>you can refer to:<br>\n<a href=\"https://arxiv.org/pdf/1912.10211.pdf\" target=\"_blank\">https://arxiv.org/pdf/1912.10211.pdf</a><br>\nPanns: Large-scale pretrained audio neural networks for audio pattern recognition</p>\n<p><img src=\"https://i.ibb.co/k9c222Y/Selection-869.png\" alt=\"https://i.ibb.co/k9c222Y/Selection-869.png\"></p>",
      "rawMarkdown": "you can refer to:\nhttps://arxiv.org/pdf/1912.10211.pdf\nPanns: Large-scale pretrained audio neural networks for audio pattern recognition\n\n![https://i.ibb.co/k9c222Y/Selection-869.png](https://i.ibb.co/k9c222Y/Selection-869.png)",
      "votes": null
    },
    {
      "id": "1512277",
      "postDate": "09/14/2021 06:30:48",
      "content": "<p>but gap between CV and LB is narrow .. with 2d cnn generally we are getting better LB wrt to CV. <br>\nany thing from this we can tell how well 1dcnn generalize wrt to 2dCNN</p>",
      "rawMarkdown": "but gap between CV and LB is narrow .. with 2d cnn generally we are getting better LB wrt to CV. \nany thing from this we can tell how well 1dcnn generalize wrt to 2dCNN",
      "votes": null
    },
    {
      "id": "1512313",
      "postDate": "09/14/2021 07:02:36",
      "content": "<p>i think the only way to is to measure the generalization gap from experiments.<br>\ni.e. try different parameters and validation sets. we need to measure difference between training and validation performances</p>",
      "rawMarkdown": "i think the only way to is to measure the generalization gap from experiments.\ni.e. try different parameters and validation sets. we need to measure difference between training and validation performances",
      "votes": null
    },
    {
      "id": "1512396",
      "postDate": "09/14/2021 08:18:53",
      "content": "<p>Yeah but how do you do transfert learning with 1d convnet ? You take pretrained model (trained on audio recognition tasks ?). My take on deep learning competition is that you always need transfert learning.</p>\n<p>Very interesting discussion. </p>",
      "rawMarkdown": "Yeah but how do you do transfert learning with 1d convnet ? You take pretrained model (trained on audio recognition tasks ?). My take on deep learning competition is that you always need transfert learning.\n\nVery interesting discussion.",
      "votes": null
    },
    {
      "id": "1512456",
      "postDate": "09/14/2021 09:11:54",
      "content": "<p>you can train from scratch<br>\n(actually, for 2d cnn , you can also do that for the smaller resnet)</p>\n<p>you can always go to github to search for trained model weights on audio, wave or GW data</p>",
      "rawMarkdown": "you can train from scratch\n(actually, for 2d cnn , you can also do that for the smaller resnet)\n\nyou can always go to github to search for trained model weights on audio, wave or GW data",
      "votes": null
    },
    {
      "id": "1512594",
      "postDate": "09/14/2021 11:45:08",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> thanks, that is very interesting analysis! if you'd like to answer bellow</p>\n<ol>\n<li><p>with k=64, padding=32 the output of 1st conv is: <code>bs, 64, 4097</code> -  do you compute paddings by hand? what follows next? --&gt; FYI Torch 1.9 accepts padding='same' </p></li>\n<li><p>what is the output of the conv stack blocks in time-dimension ? still remains 4096 or you perform any kind of pooling in between ? </p></li>\n</ol>",
      "rawMarkdown": "hengck23 thanks, that is very interesting analysis! if you'd like to answer bellow\n\n1. with k=64, padding=32 the output of 1st conv is: `bs, 64, 4097` -  do you compute paddings by hand? what follows next? --> FYI Torch 1.9 accepts padding='same' \n\n2. what is the output of the conv stack blocks in time-dimension ? still remains 4096 or you perform any kind of pooling in between ?",
      "votes": null
    },
    {
      "id": "1512963",
      "postDate": "09/14/2021 17:53:59",
      "content": "<p>good work!</p>",
      "rawMarkdown": "good work!",
      "votes": null
    },
    {
      "id": "1513174",
      "postDate": "09/14/2021 23:07:19",
      "content": "<p>i find a smart trick to ensure that the learned 1d Conv filters have both \"sine and cosine components\" like the CQT<br>\n….</p>\n<p>sift, flip your input in training and testing (TTA) instead!</p>",
      "rawMarkdown": "i find a smart trick to ensure that the learned 1d Conv filters have both \"sine and cosine components\" like the CQT\n....\n\n sift, flip your input in training and testing (TTA) instead!",
      "votes": null
    },
    {
      "id": "1514331",
      "postDate": "09/16/2021 01:28:25",
      "content": "<p>Is there any tips to train 1dCNN? I tried to follow the GW papers but my model won't learn anything. It get stuck at ROC = 0.5 forever</p>",
      "rawMarkdown": "Is there any tips to train 1dCNN? I tried to follow the GW papers but my model won't learn anything. It get stuck at ROC = 0.5 forever",
      "votes": null
    },
    {
      "id": "1515889",
      "postDate": "09/17/2021 16:07:09",
      "content": "<p>Great work!</p>",
      "rawMarkdown": "Great work!",
      "votes": null
    },
    {
      "id": "1515990",
      "postDate": "09/17/2021 18:26:09",
      "content": "<p>normalization!</p>",
      "rawMarkdown": "normalization!",
      "votes": null
    },
    {
      "id": "1517001",
      "postDate": "09/19/2021 04:50:15",
      "content": "<p>Tnx for answering. Do you mind elaborating a bit more? or maybe pointing me to some good material? Thanks</p>",
      "rawMarkdown": "Tnx for answering. Do you mind elaborating a bit more? or maybe pointing me to some good material? Thanks",
      "votes": null
    },
    {
      "id": "1559723",
      "postDate": "10/27/2021 07:09:30",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1509905,
      "author_name": "zarif98sjs",
      "author_url": "",
      "post_date": "09/11/2021 19:13:51",
      "content": "<blockquote>\n  <p>The CQT is not used in the CNN</p>\n</blockquote>\n<p>What's used ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1509910,
          "author_name": "brachester",
          "author_url": "",
          "post_date": "09/11/2021 19:17:07",
          "content": "<p>It's a 1d cnn, you're thinking of 2d cnn. In this case, he probably used the raw time series.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1509912,
          "author_name": "zarif98sjs",
          "author_url": "",
          "post_date": "09/11/2021 19:19:16",
          "content": "<p>Oops sorry, missed that part !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1510028,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "09/12/2021 01:09:36",
      "content": "<p>update1.<br>\n(selection of max kernel (red box) could be wrong as I did not consider the weights of the next bn layer)</p>\n<p><img src=\"https://i.ibb.co/n3gQzy8/Selection-851.png\" alt=\"https://i.ibb.co/n3gQzy8/Selection-851.png\"><br>\n<img src=\"https://i.ibb.co/C70jXPQ/Selection-850.png\" alt=\"https://i.ibb.co/C70jXPQ/Selection-850.png\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1510036,
          "author_name": "authman",
          "author_url": "",
          "post_date": "09/12/2021 02:11:13",
          "content": "<p>Interesting. It seems as though you're running your 1d cnn on frequency domain magnitudes as opposed to time domain?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1510038,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "09/12/2021 02:16:53",
          "content": "<p>no. i am running on time domain</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1510029,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "09/12/2021 01:17:44",
      "content": "<p>if you are interested in how to design waveform 1d CNN and compare it with spectrogram 2d CNN, refer to:</p>\n<p>ISMIR 2019 tutorial: waveform-based music processing with deep learning<br>\n<a href=\"https://zenodo.org/record/3529714#.YT1UZnUzbCJ\" target=\"_blank\">https://zenodo.org/record/3529714#.YT1UZnUzbCJ</a></p>\n<p>Learning Multiscale Features Directly From Waveforms<br>\n<a href=\"http://research.baidu.com/Public/uploads/5ac04fa6a3f28.PDF\" target=\"_blank\">http://research.baidu.com/Public/uploads/5ac04fa6a3f28.PDF</a></p>\n<p>SampleCNN: End-to-End Deep Convolutional Neural Networks Using Very Small Filters for Music Classification</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1511136,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "09/13/2021 06:43:39",
      "content": "<p>interesting, whitening don't work for both 1d cnn and 2d cnn.<br>\n(maybe 2 sec is too short to estimate noise power)</p>\n<p>i transfer the 2d-cnn signal processing for 1d-cnn, the performance of 1d-cnn is greatly increased</p>\n<p>simple 8 layer 1dcnn has cv = 0.868902<br>\n(45 min to train up to 6 epoach)</p>\n<p>the same wave signal when transform with CQT(128x256)+eff-b0 gives cv = 0.8716</p>\n<p>ensembling gives  cv =0.873843</p>\n<hr>\n<p>bandpass seems to be a must for 1d-cnn</p>",
      "votes": null,
      "replies": [
        {
          "id": 1511166,
          "author_name": "pheadrus",
          "author_url": "",
          "post_date": "09/13/2021 07:15:39",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Are you saying you are feeding raw waveform + bandpass to 1D CNNs? Any other pre-processing that you can tell us about?  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1511309,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "09/13/2021 10:11:16",
          "content": "<p>after increasing the parameters of the 1d-cnn:<br>\n1d-cnn auc  cv = 0.870425<br>\nensemble auc cv = 0.874456</p>\n<p></p>\n<p>(this is wrong. i forget that the y-axis is log scale)</p>\n<p><img src=\"https://i.ibb.co/xhFrMYs/Selection-864.png\" alt=\"https://i.ibb.co/xhFrMYs/Selection-864.png\"><br>\n<img src=\"https://i.ibb.co/fCX7YLw/Selection-863.png\" alt=\"https://i.ibb.co/fCX7YLw/Selection-863.png\"></p>\n<hr>\n<p>instead of ensembling, why not train jointly (or partially jointly, e.g. freeze one of the network if you have limited GPU memory)?</p>\n<p>if the 1d-CNN of the ensemble gets accurate enough, one might want to try distillation</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1511470,
          "author_name": "authman",
          "author_url": "",
          "post_date": "09/13/2021 12:58:52",
          "content": "<blockquote>\n  <p>instead of ensembling, why not train jointly</p>\n</blockquote>\n<p>I did this. But 1d cnn learning schedule is different from 2d, so either separate training then joining needs to be done. Or for more optimal results, careful finetuning of differential learning rates, which I haven't explored.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1511500,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "09/13/2021 13:27:41",
          "content": "<p>i just use trained 1d and 2d network as feature extractors.<br>\nthey are freeze when training a combined network.</p>\n<p>but even so, there seems to be overfitting, so maybe it is not easy to train jointly</p>\n<p>we can flip left-right and up-down for noise wave form. can we do that for GW samples as well?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1511550,
          "author_name": "martynoveduard",
          "author_url": "",
          "post_date": "09/13/2021 14:09:41",
          "content": "<p>I'm not sure that effect of reduction of FP comes from exactly \"ensemble of 1D-net and 2D-net\" but instead because of the nature of ensemble itself<br>\nWhat I'm saying is that you could achieve same result with ensembling using only 2D-nets, I suppose<br>\nIt would be interesting to look on the same plots, but with ensemble of 2d-nets only, I should try to do that myself I suppose</p>\n<p>By the way, how to you define \"pos\" or \"neg\" in your plots? <br>\nBecause usually we compute ROC-AUC on the raw probabilities, but judging from the second plot you divided samples into \"pos/neg\" somehow.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1511551,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "09/13/2021 14:11:56",
          "content": "<p>By the way, how to you define \"pos\" or \"neg\" in your plots?<br>\n…</p>\n<p>we have the ground truth label in validation for the distribution plots.<br>\nfor ROC, it is for all train samples (i.e. pos and neg)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1511565,
          "author_name": "authman",
          "author_url": "",
          "post_date": "09/13/2021 14:19:09",
          "content": "<p><a href=\"https://www.kaggle.com/rimuru\" target=\"_blank\">@rimuru</a> try something like:</p>\n<pre><code>hist, edges = np.histogram(z.probs, bins=100)\nplt.plot(edges[:-1], hist, linewidth=1)\nplt.yscale('log')\nplt.show()\n</code></pre>\n<p>But split z by train target or oof target.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1511811,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "09/13/2021 17:52:39",
          "content": "<p>after more augmentations:</p>\n<p></p>\n<p>for one fold only:</p>\n<p>2d-cnn ( CQT 128x256 + eff-b0)  auc cv = 0.872293 <br>\n1d-cnn (wave input to 8 layer temporal conv) auc cv = 0.873005 (lb =0.8746)<br>\nensemble:  0.875676 (this is almost the same results as effnet b7)</p>\n<p>conclusion:</p>\n<ol>\n<li>it is possible that 1d-cnn outperforms 2d-cnn</li>\n<li>you can get good results using low computation model</li>\n<li>it is possible to train from scratch</li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1511850,
          "author_name": "amedprof",
          "author_url": "",
          "post_date": "09/13/2021 18:18:45",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> are you using the 1d-cnn on the signal directly or on the spectrogram ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1511860,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "09/13/2021 18:23:32",
          "content": "<pre><code>class Net(nn.Module):\n    # 1D convolutional neural network.   \n    def __init__(self):\n        super().__init__()\n        self.cnn1 =  nnn.Conv1d(3, 64, kernel_size=64, padding=32)\n       ... etc ...\n\n\n    def forward(self, wave):\n        x = self.cnn1(wave)  \n       ... etc ...\n\n\ndef run_check_net():\n    batch_size = 5\n    wave = torch.randn(batch_size, 3, 4096) \n\n    net = Net() \n    signal= net(wave)\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1512153,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "09/14/2021 03:21:05",
          "content": "<p><img src=\"https://i.ibb.co/5F3dvjD/Selection-868.png\" alt=\"https://i.ibb.co/5F3dvjD/Selection-868.png\"></p>\n<p>my guess.</p>\n<p>do note that performance is limited by data.<br>\nanything that is in the \"front\" (e.g. data, augmentation, signal processing) set the limit of performance.<br>\nanything in the \"back\" determines how close you can reach the limit</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1512264,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "09/14/2021 06:20:48",
          "content": "<p>you can refer to:<br>\n<a href=\"https://arxiv.org/pdf/1912.10211.pdf\" target=\"_blank\">https://arxiv.org/pdf/1912.10211.pdf</a><br>\nPanns: Large-scale pretrained audio neural networks for audio pattern recognition</p>\n<p><img src=\"https://i.ibb.co/k9c222Y/Selection-869.png\" alt=\"https://i.ibb.co/k9c222Y/Selection-869.png\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1512277,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "09/14/2021 06:30:48",
          "content": "<p>but gap between CV and LB is narrow .. with 2d cnn generally we are getting better LB wrt to CV. <br>\nany thing from this we can tell how well 1dcnn generalize wrt to 2dCNN</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1512313,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "09/14/2021 07:02:36",
          "content": "<p>i think the only way to is to measure the generalization gap from experiments.<br>\ni.e. try different parameters and validation sets. we need to measure difference between training and validation performances</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1512594,
          "author_name": "imeintanis",
          "author_url": "",
          "post_date": "09/14/2021 11:45:08",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> thanks, that is very interesting analysis! if you'd like to answer bellow</p>\n<ol>\n<li><p>with k=64, padding=32 the output of 1st conv is: <code>bs, 64, 4097</code> -  do you compute paddings by hand? what follows next? --&gt; FYI Torch 1.9 accepts padding='same' </p></li>\n<li><p>what is the output of the conv stack blocks in time-dimension ? still remains 4096 or you perform any kind of pooling in between ? </p></li>\n</ol>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1512396,
      "author_name": "forbu1411",
      "author_url": "",
      "post_date": "09/14/2021 08:18:53",
      "content": "<p>Yeah but how do you do transfert learning with 1d convnet ? You take pretrained model (trained on audio recognition tasks ?). My take on deep learning competition is that you always need transfert learning.</p>\n<p>Very interesting discussion. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1512456,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "09/14/2021 09:11:54",
          "content": "<p>you can train from scratch<br>\n(actually, for 2d cnn , you can also do that for the smaller resnet)</p>\n<p>you can always go to github to search for trained model weights on audio, wave or GW data</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1512963,
      "author_name": "faaizhashmi",
      "author_url": "",
      "post_date": "09/14/2021 17:53:59",
      "content": "<p>good work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1513174,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "09/14/2021 23:07:19",
      "content": "<p>i find a smart trick to ensure that the learned 1d Conv filters have both \"sine and cosine components\" like the CQT<br>\n….</p>\n<p>sift, flip your input in training and testing (TTA) instead!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1514331,
      "author_name": "coldfir3",
      "author_url": "",
      "post_date": "09/16/2021 01:28:25",
      "content": "<p>Is there any tips to train 1dCNN? I tried to follow the GW papers but my model won't learn anything. It get stuck at ROC = 0.5 forever</p>",
      "votes": null,
      "replies": [
        {
          "id": 1515990,
          "author_name": "abhishek",
          "author_url": "",
          "post_date": "09/17/2021 18:26:09",
          "content": "<p>normalization!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1517001,
          "author_name": "coldfir3",
          "author_url": "",
          "post_date": "09/19/2021 04:50:15",
          "content": "<p>Tnx for answering. Do you mind elaborating a bit more? or maybe pointing me to some good material? Thanks</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1515889,
      "author_name": "katew19",
      "author_url": "",
      "post_date": "09/17/2021 16:07:09",
      "content": "<p>Great work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1559723,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 07:09:30",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1509426": "this is a preview only.\nmore analysis and results coming up soon!\n\n![https://i.ibb.co/7Cm6Fwy/Selection-839.png](https://i.ibb.co/7Cm6Fwy/Selection-839.png)",
    "1509905": "> The CQT is not used in the CNN\n\nWhat's used ?",
    "1509910": "It's a 1d cnn, you're thinking of 2d cnn. In this case, he probably used the raw time series.",
    "1509912": "Oops sorry, missed that part !",
    "1510028": "update1.\n(selection of max kernel (red box) could be wrong as I did not consider the weights of the next bn layer)\n\n![https://i.ibb.co/n3gQzy8/Selection-851.png](https://i.ibb.co/n3gQzy8/Selection-851.png)\n![https://i.ibb.co/C70jXPQ/Selection-850.png](https://i.ibb.co/C70jXPQ/Selection-850.png)",
    "1510029": "if you are interested in how to design waveform 1d CNN and compare it with spectrogram 2d CNN, refer to:\n\nISMIR 2019 tutorial: waveform-based music processing with deep learning\nhttps://zenodo.org/record/3529714#.YT1UZnUzbCJ\n\nLearning Multiscale Features Directly From Waveforms\nhttp://research.baidu.com/Public/uploads/5ac04fa6a3f28.PDF\n\n\nSampleCNN: End-to-End Deep Convolutional Neural Networks Using Very Small Filters for Music Classification",
    "1510036": "Interesting. It seems as though you're running your 1d cnn on frequency domain magnitudes as opposed to time domain?",
    "1510038": "no. i am running on time domain",
    "1511136": "interesting, whitening don't work for both 1d cnn and 2d cnn.\n(maybe 2 sec is too short to estimate noise power)\n\ni transfer the 2d-cnn signal processing for 1d-cnn, the performance of 1d-cnn is greatly increased\n\nsimple 8 layer 1dcnn has cv = 0.868902\n(45 min to train up to 6 epoach)\n\nthe same wave signal when transform with CQT(128x256)+eff-b0 gives cv = 0.8716\n\nensembling gives  cv =0.873843\n\n---\n\nbandpass seems to be a must for 1d-cnn",
    "1511166": "hengck23 Are you saying you are feeding raw waveform + bandpass to 1D CNNs? Any other pre-processing that you can tell us about?",
    "1511309": "after increasing the parameters of the 1d-cnn:\n1d-cnn auc  cv = 0.870425\nensemble auc cv = 0.874456\n\n\n~~where did the improvement come from?\ni think we did not detect more GW. But we did reduce the number of FP~~\n\n(this is wrong. i forget that the y-axis is log scale)\n\n\n![https://i.ibb.co/xhFrMYs/Selection-864.png](https://i.ibb.co/xhFrMYs/Selection-864.png)\n![https://i.ibb.co/fCX7YLw/Selection-863.png](https://i.ibb.co/fCX7YLw/Selection-863.png)\n\n---\ninstead of ensembling, why not train jointly (or partially jointly, e.g. freeze one of the network if you have limited GPU memory)?\n\nif the 1d-CNN of the ensemble gets accurate enough, one might want to try distillation",
    "1511470": "> instead of ensembling, why not train jointly\n\nI did this. But 1d cnn learning schedule is different from 2d, so either separate training then joining needs to be done. Or for more optimal results, careful finetuning of differential learning rates, which I haven't explored.",
    "1511500": "i just use trained 1d and 2d network as feature extractors.\nthey are freeze when training a combined network.\n\nbut even so, there seems to be overfitting, so maybe it is not easy to train jointly\n\n\nwe can flip left-right and up-down for noise wave form. can we do that for GW samples as well?",
    "1511550": "I'm not sure that effect of reduction of FP comes from exactly \"ensemble of 1D-net and 2D-net\" but instead because of the nature of ensemble itself\nWhat I'm saying is that you could achieve same result with ensembling using only 2D-nets, I suppose\nIt would be interesting to look on the same plots, but with ensemble of 2d-nets only, I should try to do that myself I suppose\n\nBy the way, how to you define \"pos\" or \"neg\" in your plots? \nBecause usually we compute ROC-AUC on the raw probabilities, but judging from the second plot you divided samples into \"pos/neg\" somehow.",
    "1511551": "By the way, how to you define \"pos\" or \"neg\" in your plots?\n...\n\nwe have the ground truth label in validation for the distribution plots.\nfor ROC, it is for all train samples (i.e. pos and neg)",
    "1511565": "rimuru try something like:\n\n```\nhist, edges = np.histogram(z.probs, bins=100)\nplt.plot(edges[:-1], hist, linewidth=1)\nplt.yscale('log')\nplt.show()\n```\n\nBut split z by train target or oof target.",
    "1511811": "after more augmentations:\n\n~~1d-cnn auc cv = 0.872519\nensemble auc cv = 0.875179  (this is almost the same results as effnet b7)~~\n\nfor one fold only:\n\n2d-cnn ( CQT 128x256 + eff-b0)  auc cv = 0.872293 \n1d-cnn (wave input to 8 layer temporal conv) auc cv = 0.873005 (lb =0.8746)\nensemble:  0.875676 (this is almost the same results as effnet b7)\n\n\nconclusion:\n1. it is possible that 1d-cnn outperforms 2d-cnn\n2. you can get good results using low computation model\n3. it is possible to train from scratch",
    "1511850": "hengck23 are you using the 1d-cnn on the signal directly or on the spectrogram ?",
    "1511860": "```\nclass Net(nn.Module):\n    # 1D convolutional neural network.   \n    def __init__(self):\n        super().__init__()\n        self.cnn1 =  nnn.Conv1d(3, 64, kernel_size=64, padding=32)\n       ... etc ...\n     \n\n    def forward(self, wave):\n        x = self.cnn1(wave)  \n       ... etc ...\n\n\ndef run_check_net():\n    batch_size = 5\n    wave = torch.randn(batch_size, 3, 4096) \n\n    net = Net() \n    signal= net(wave)\n\n```",
    "1512153": "![https://i.ibb.co/5F3dvjD/Selection-868.png](https://i.ibb.co/5F3dvjD/Selection-868.png)\n\nmy guess.\n\ndo note that performance is limited by data.\nanything that is in the \"front\" (e.g. data, augmentation, signal processing) set the limit of performance.\nanything in the \"back\" determines how close you can reach the limit",
    "1512264": "you can refer to:\nhttps://arxiv.org/pdf/1912.10211.pdf\nPanns: Large-scale pretrained audio neural networks for audio pattern recognition\n\n![https://i.ibb.co/k9c222Y/Selection-869.png](https://i.ibb.co/k9c222Y/Selection-869.png)",
    "1512277": "but gap between CV and LB is narrow .. with 2d cnn generally we are getting better LB wrt to CV. \nany thing from this we can tell how well 1dcnn generalize wrt to 2dCNN",
    "1512313": "i think the only way to is to measure the generalization gap from experiments.\ni.e. try different parameters and validation sets. we need to measure difference between training and validation performances",
    "1512396": "Yeah but how do you do transfert learning with 1d convnet ? You take pretrained model (trained on audio recognition tasks ?). My take on deep learning competition is that you always need transfert learning.\n\nVery interesting discussion.",
    "1512456": "you can train from scratch\n(actually, for 2d cnn , you can also do that for the smaller resnet)\n\nyou can always go to github to search for trained model weights on audio, wave or GW data",
    "1512594": "hengck23 thanks, that is very interesting analysis! if you'd like to answer bellow\n\n1. with k=64, padding=32 the output of 1st conv is: `bs, 64, 4097` -  do you compute paddings by hand? what follows next? --> FYI Torch 1.9 accepts padding='same' \n\n2. what is the output of the conv stack blocks in time-dimension ? still remains 4096 or you perform any kind of pooling in between ?",
    "1512963": "good work!",
    "1513174": "i find a smart trick to ensure that the learned 1d Conv filters have both \"sine and cosine components\" like the CQT\n....\n\n sift, flip your input in training and testing (TTA) instead!",
    "1514331": "Is there any tips to train 1dCNN? I tried to follow the GW papers but my model won't learn anything. It get stuck at ROC = 0.5 forever",
    "1515889": "Great work!",
    "1515990": "normalization!",
    "1517001": "Tnx for answering. Do you mind elaborating a bit more? or maybe pointing me to some good material? Thanks",
    "1559723": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris"
  },
  "source": "meta"
}