{
  "id": 511763,
  "title": "What is the intuition behind feeding signal to vision model? ",
  "url": "/competitions/birdclef-2024/discussion/511763",
  "author_name": "coolz",
  "post_date": "2024-06-12T02:30:02.051000",
  "votes": 16,
  "comment_count": 8,
  "views": 0,
  "content": "<p>First of all,  i am very happy to become a GM. In my last 2 competitions, i use the same methods, using vision model to deal with 1d signal, it helps a lot. Here, i introduce the method in a simple way.  </p>\n<p>In my thought, if the spectrum works well, there is no reason that raw signal cannot work !  </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1562037%2F787b347daff61fdf8fbdda3cefd791a7%2Fimg.png?generation=1718160172309414&amp;alt=media\"><br>\n                                        Figure1</p>\n<p>In figure1 the 0-15 is the raw signal array. After reshape, it becomes to the metric array.</p>\n<p>A 3x3 conv kernel actually equal to a special conv1d that works on the raw signal( the red number).By stacking the conv2d layer, we can get the feature.</p>\n<p>Additional, with many padding operators in vision model, we can get the position info better.And in some degree, that conv1d works not generally well,is because the position information cannot captured very well.With a proper reshape,with padding conv2d can do that well.(I din't sure for this conclusion, just intuition)</p>\n<p>However we don't have very good pretrained 1d cnn that can work pretty well with 1d raw signal. In this way, we have plenty of pretrained models. </p>\n<p>Enjoy it.  </p>",
  "messages": [
    {
      "id": 2867673,
      "postDate": "2024-06-12T02:30:02.050Z",
      "content": "<p>First of all,  i am very happy to become a GM. In my last 2 competitions, i use the same methods, using vision model to deal with 1d signal, it helps a lot. Here, i introduce the method in a simple way.  </p>\n<p>In my thought, if the spectrum works well, there is no reason that raw signal cannot work !  </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1562037%2F787b347daff61fdf8fbdda3cefd791a7%2Fimg.png?generation=1718160172309414&amp;alt=media\"><br>\n                                        Figure1</p>\n<p>In figure1 the 0-15 is the raw signal array. After reshape, it becomes to the metric array.</p>\n<p>A 3x3 conv kernel actually equal to a special conv1d that works on the raw signal( the red number).By stacking the conv2d layer, we can get the feature.</p>\n<p>Additional, with many padding operators in vision model, we can get the position info better.And in some degree, that conv1d works not generally well,is because the position information cannot captured very well.With a proper reshape,with padding conv2d can do that well.(I din't sure for this conclusion, just intuition)</p>\n<p>However we don't have very good pretrained 1d cnn that can work pretty well with 1d raw signal. In this way, we have plenty of pretrained models. </p>\n<p>Enjoy it.  </p>",
      "rawMarkdown": "First of all,  i am very happy to become a GM. In my last 2 competitions, i use the same methods, using vision model to deal with 1d signal, it helps a lot. Here, i introduce the method in a simple way.  \n\nIn my thought, if the spectrum works well, there is no reason that raw signal cannot work !  \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1562037%2F787b347daff61fdf8fbdda3cefd791a7%2Fimg.png?generation=1718160172309414&alt=media)\n                                        Figure1\n\nIn figure1 the 0-15 is the raw signal array. After reshape, it becomes to the metric array.\n\n  \nA 3x3 conv kernel actually equal to a special conv1d that works on the raw signal( the red number).By stacking the conv2d layer, we can get the feature.\n\n  \nAdditional, with many padding operators in vision model, we can get the position info better.And in some degree, that conv1d works not generally well,is because the position information cannot captured very well.With a proper reshape,with padding conv2d can do that well.(I din't sure for this conclusion, just intuition)\n\n\nHowever we don't have very good pretrained 1d cnn that can work pretty well with 1d raw signal. In this way, we have plenty of pretrained models. \n\nEnjoy it.  \n\n\n",
      "votes": 16
    },
    {
      "id": 2867680,
      "postDate": "2024-06-12T02:45:32.337Z",
      "content": "<p>Could it be due to domain shift? During my experiments in this competition, I found that reducing domain differences was challenging when using spectral/image-based 2D methods. And the results were even highly sensitive to changes in hop_size/image_size. So i think the domain shift in raw signal maybe small than spectral.</p>",
      "rawMarkdown": "Could it be due to domain shift? During my experiments in this competition, I found that reducing domain differences was challenging when using spectral/image-based 2D methods. And the results were even highly sensitive to changes in hop_size/image_size. So i think the domain shift in raw signal maybe small than spectral.",
      "votes": 1,
      "replies": [
        {
          "id": 2867683,
          "postDate": "2024-06-12T02:55:36.580Z",
          "content": "<p>I don't think so. Refer to the score, all of us didn't solve this problem very well, including the top. I am upset by this.</p>",
          "rawMarkdown": "I don't think so. Refer to the score, all of us didn't solve this problem very well, including the top. I am upset by this.",
          "votes": 3,
          "replies": [
            {
              "id": 2867693,
              "postDate": "2024-06-12T03:09:40.080Z",
              "content": "<p>This is truly a great idea. We can consider the x-axis to represent local or slice-wise information, while the y-axis represents the aggregation of global information through downsampling, pooling, or interpolation. Congratulations on becoming a GM!</p>",
              "rawMarkdown": "This is truly a great idea. We can consider the x-axis to represent local or slice-wise information, while the y-axis represents the aggregation of global information through downsampling, pooling, or interpolation. Congratulations on becoming a GM!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2877779,
      "postDate": "2024-06-18T15:04:39.680Z",
      "content": "<p>Hi, you might be interested in HiFiGAN, which used multi-resolution waveforms in the discriminator to improve audio generation:<br>\n<a href=\"https://arxiv.org/abs/2010.05646\" target=\"_blank\">https://arxiv.org/abs/2010.05646</a></p>\n<p>I believe they used 1d convolutions on the waveforms, in addition to spectrogram classifiers. As a GAN, these many classifiers provided signal for improvement for the main generative model.</p>",
      "rawMarkdown": "Hi, you might be interested in HiFiGAN, which used multi-resolution waveforms in the discriminator to improve audio generation:\nhttps://arxiv.org/abs/2010.05646\n\nI believe they used 1d convolutions on the waveforms, in addition to spectrogram classifiers. As a GAN, these many classifiers provided signal for improvement for the main generative model.",
      "votes": 2,
      "replies": [
        {
          "id": 2880144,
          "postDate": "2024-06-20T04:38:43.317Z",
          "content": "<p>Yeah.Thanks for the paper. The Multi-Period Discriminator(MPD) is quite similar, and it's still use some kind of 1d conv( conv2d with dim=1 in one direction). </p>",
          "rawMarkdown": "Yeah.Thanks for the paper. The Multi-Period Discriminator(MPD) is quite similar, and it's still use some kind of 1d conv( conv2d with dim=1 in one direction). "
        }
      ]
    },
    {
      "id": 2872180,
      "postDate": "2024-06-14T16:13:15.723Z",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/cooolz\" target=\"_blank\">@cooolz</a>, thanks for sharing this great idea. Do you tune the height and width when stacking the signal into a metric array?</p>\n<p>For example, I created two cases below. I think example A acts like a 1D-conv with dilation=4, and then example B acts like a 1D-conv with a dilation=8. How do you find the best shape for the signal?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2Fee263437450ba6e6be0bddc47149c206%2FUntitled%20Diagram%20(1).jpg?generation=1718381331380110&amp;alt=media\"></p>",
      "rawMarkdown": "Hey @cooolz, thanks for sharing this great idea. Do you tune the height and width when stacking the signal into a metric array?\n\nFor example, I created two cases below. I think example A acts like a 1D-conv with dilation=4, and then example B acts like a 1D-conv with a dilation=8. How do you find the best shape for the signal?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2Fee263437450ba6e6be0bddc47149c206%2FUntitled%20Diagram%20(1).jpg?generation=1718381331380110&alt=media)",
      "votes": -1,
      "replies": [
        {
          "id": 2880106,
          "postDate": "2024-06-20T04:10:08.687Z",
          "content": "<p>For me, it's experiment depends.  </p>\n<p>But i thought there might be a better way to do the packing. I might going to explore more about the method in some new competition.</p>\n<p>Yes it's like dilation conv1d. But conv2d has more padding operators in x and y directions. And it's not that big in size.  </p>\n<p>For conv1d, though it's padded, the sequence is too long, it's difficult to capture location information. </p>\n<p>And there is pretrained weights with conv2d.  Here is what i thought.</p>",
          "rawMarkdown": "For me, it's experiment depends.  \n\nBut i thought there might be a better way to do the packing. I might going to explore more about the method in some new competition.\n\nYes it's like dilation conv1d. But conv2d has more padding operators in x and y directions. And it's not that big in size.  \n\nFor conv1d, though it's padded, the sequence is too long, it's difficult to capture location information. \n\nAnd there is pretrained weights with conv2d.  Here is what i thought.",
          "votes": 2,
          "replies": [
            {
              "id": 2880116,
              "postDate": "2024-06-20T04:15:00.657Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2867680,
      "author_name": "tanxxx",
      "author_url": "",
      "post_date": "2024-06-12T02:45:32.337000",
      "content": "<p>Could it be due to domain shift? During my experiments in this competition, I found that reducing domain differences was challenging when using spectral/image-based 2D methods. And the results were even highly sensitive to changes in hop_size/image_size. So i think the domain shift in raw signal maybe small than spectral.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2867683,
          "author_name": "coolz",
          "author_url": "",
          "post_date": "2024-06-12T02:55:36.580000",
          "content": "<p>I don't think so. Refer to the score, all of us didn't solve this problem very well, including the top. I am upset by this.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2867693,
              "author_name": "tanxxx",
              "author_url": "",
              "post_date": "2024-06-12T03:09:40.080000",
              "content": "<p>This is truly a great idea. We can consider the x-axis to represent local or slice-wise information, while the y-axis represents the aggregation of global information through downsampling, pooling, or interpolation. Congratulations on becoming a GM!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2877779,
      "author_name": "Tom Denton",
      "author_url": "",
      "post_date": "2024-06-18T15:04:39.680000",
      "content": "<p>Hi, you might be interested in HiFiGAN, which used multi-resolution waveforms in the discriminator to improve audio generation:<br>\n<a href=\"https://arxiv.org/abs/2010.05646\" target=\"_blank\">https://arxiv.org/abs/2010.05646</a></p>\n<p>I believe they used 1d convolutions on the waveforms, in addition to spectrogram classifiers. As a GAN, these many classifiers provided signal for improvement for the main generative model.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2880144,
          "author_name": "coolz",
          "author_url": "",
          "post_date": "2024-06-20T04:38:43.317000",
          "content": "<p>Yeah.Thanks for the paper. The Multi-Period Discriminator(MPD) is quite similar, and it's still use some kind of 1d conv( conv2d with dim=1 in one direction). </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2872180,
      "author_name": "Bartley",
      "author_url": "",
      "post_date": "2024-06-14T16:13:15.723000",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/cooolz\" target=\"_blank\">@cooolz</a>, thanks for sharing this great idea. Do you tune the height and width when stacking the signal into a metric array?</p>\n<p>For example, I created two cases below. I think example A acts like a 1D-conv with dilation=4, and then example B acts like a 1D-conv with a dilation=8. How do you find the best shape for the signal?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2Fee263437450ba6e6be0bddc47149c206%2FUntitled%20Diagram%20(1).jpg?generation=1718381331380110&amp;alt=media\"></p>",
      "votes": -1,
      "replies": [
        {
          "id": 2880106,
          "author_name": "coolz",
          "author_url": "",
          "post_date": "2024-06-20T04:10:08.687000",
          "content": "<p>For me, it's experiment depends.  </p>\n<p>But i thought there might be a better way to do the packing. I might going to explore more about the method in some new competition.</p>\n<p>Yes it's like dilation conv1d. But conv2d has more padding operators in x and y directions. And it's not that big in size.  </p>\n<p>For conv1d, though it's padded, the sequence is too long, it's difficult to capture location information. </p>\n<p>And there is pretrained weights with conv2d.  Here is what i thought.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2880116,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-06-20T04:15:00.657000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2867673": "First of all,  i am very happy to become a GM. In my last 2 competitions, i use the same methods, using vision model to deal with 1d signal, it helps a lot. Here, i introduce the method in a simple way.  \n\nIn my thought, if the spectrum works well, there is no reason that raw signal cannot work !  \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1562037%2F787b347daff61fdf8fbdda3cefd791a7%2Fimg.png?generation=1718160172309414&alt=media)\n                                        Figure1\n\nIn figure1 the 0-15 is the raw signal array. After reshape, it becomes to the metric array.\n\n  \nA 3x3 conv kernel actually equal to a special conv1d that works on the raw signal( the red number).By stacking the conv2d layer, we can get the feature.\n\n  \nAdditional, with many padding operators in vision model, we can get the position info better.And in some degree, that conv1d works not generally well,is because the position information cannot captured very well.With a proper reshape,with padding conv2d can do that well.(I din't sure for this conclusion, just intuition)\n\n\nHowever we don't have very good pretrained 1d cnn that can work pretty well with 1d raw signal. In this way, we have plenty of pretrained models. \n\nEnjoy it.  \n\n\n",
    "2867680": "Could it be due to domain shift? During my experiments in this competition, I found that reducing domain differences was challenging when using spectral/image-based 2D methods. And the results were even highly sensitive to changes in hop_size/image_size. So i think the domain shift in raw signal maybe small than spectral.",
    "2877779": "Hi, you might be interested in HiFiGAN, which used multi-resolution waveforms in the discriminator to improve audio generation:\nhttps://arxiv.org/abs/2010.05646\n\nI believe they used 1d convolutions on the waveforms, in addition to spectrogram classifiers. As a GAN, these many classifiers provided signal for improvement for the main generative model.",
    "2872180": "Hey @cooolz, thanks for sharing this great idea. Do you tune the height and width when stacking the signal into a metric array?\n\nFor example, I created two cases below. I think example A acts like a 1D-conv with dilation=4, and then example B acts like a 1D-conv with a dilation=8. How do you find the best shape for the signal?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2Fee263437450ba6e6be0bddc47149c206%2FUntitled%20Diagram%20(1).jpg?generation=1718381331380110&alt=media)"
  }
}