{
  "id": 502574,
  "title": "Why does image spectrogram input work better than a 1d conv model?",
  "url": "/competitions/birdclef-2024/discussion/502574",
  "author_name": "",
  "post_date": "2024-05-14T00:35:33.160530500Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>This is my first BIRDCLEF competition, however, I do not understand why image data performs better than a simple 1d convolution time-series model. Theoretically in image convolution the quality of the notes will be lost when downscaling 10/12/16 bit ADC volume values to 256 or 8 bit colors is this true or am I missing something? Many thanks for answering my question!</p>",
  "messages": [
    {
      "id": "2811850",
      "postDate": "05/14/2024 00:35:33",
      "content": "<p>This is my first BIRDCLEF competition, however, I do not understand why image data performs better than a simple 1d convolution time-series model. Theoretically in image convolution the quality of the notes will be lost when downscaling 10/12/16 bit ADC volume values to 256 or 8 bit colors is this true or am I missing something? Many thanks for answering my question!</p>",
      "rawMarkdown": "This is my first BIRDCLEF competition, however, I do not understand why image data performs better than a simple 1d convolution time-series model. Theoretically in image convolution the quality of the notes will be lost when downscaling 10/12/16 bit ADC volume values to 256 or 8 bit colors is this true or am I missing something? Many thanks for answering my question!",
      "votes": null
    },
    {
      "id": "2811956",
      "postDate": "05/14/2024 02:49:21",
      "content": "<p>I am a speech signal processing engineer, and from my point,<br>\n1d time-series signal indeed have much information, but compared with 2d signals, 1d signal is too abstract, thus, for human, directly analyse with it is very challenging. On the other hand, if you have trained a 1d convolution time-series model before, you might find that 1d models are always suffered with overfitting problems and large number of parameters (if with large convolutional kernel ), so if the data number is limited, 2d model may be my first choice.</p>\n<p>\"quality of the notes will be lost when downscaling 10/12/16 bit ADC volume values to 256 or 8 bit colors \"<br>\nThere is indeed a loss in quality during downsampling. For instance, if we only consider audio signals with a 16-bit resolution, downsampling or normalizing to 16 bits would involve dividing the audio signal by 2**16, and similarly, downsampling or normalizing to 8 bits would involve dividing the audio signal by 2 **8. Thus, the difference between these two downsampling methods is merely a scalar factor, which in audio processing terms, can be regarded as a change in volume.</p>",
      "rawMarkdown": "I am a speech signal processing engineer, and from my point,\n1d time-series signal indeed have much information, but compared with 2d signals, 1d signal is too abstract, thus, for human, directly analyse with it is very challenging. On the other hand, if you have trained a 1d convolution time-series model before, you might find that 1d models are always suffered with overfitting problems and large number of parameters (if with large convolutional kernel ), so if the data number is limited, 2d model may be my first choice.\n\n\"quality of the notes will be lost when downscaling 10/12/16 bit ADC volume values to 256 or 8 bit colors \"\nThere is indeed a loss in quality during downsampling. For instance, if we only consider audio signals with a 16-bit resolution, downsampling or normalizing to 16 bits would involve dividing the audio signal by 2**16, and similarly, downsampling or normalizing to 8 bits would involve dividing the audio signal by 2 **8. Thus, the difference between these two downsampling methods is merely a scalar factor, which in audio processing terms, can be regarded as a change in volume.",
      "votes": null
    },
    {
      "id": "2811958",
      "postDate": "05/14/2024 02:58:52",
      "content": "<p>Thank you very much I was having trouble understanding this one key point! 😁</p>",
      "rawMarkdown": "Thank you very much I was having trouble understanding this one key point! 😁",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2811956,
      "author_name": "tanxxx",
      "author_url": "",
      "post_date": "05/14/2024 02:49:21",
      "content": "<p>I am a speech signal processing engineer, and from my point,<br>\n1d time-series signal indeed have much information, but compared with 2d signals, 1d signal is too abstract, thus, for human, directly analyse with it is very challenging. On the other hand, if you have trained a 1d convolution time-series model before, you might find that 1d models are always suffered with overfitting problems and large number of parameters (if with large convolutional kernel ), so if the data number is limited, 2d model may be my first choice.</p>\n<p>\"quality of the notes will be lost when downscaling 10/12/16 bit ADC volume values to 256 or 8 bit colors \"<br>\nThere is indeed a loss in quality during downsampling. For instance, if we only consider audio signals with a 16-bit resolution, downsampling or normalizing to 16 bits would involve dividing the audio signal by 2**16, and similarly, downsampling or normalizing to 8 bits would involve dividing the audio signal by 2 **8. Thus, the difference between these two downsampling methods is merely a scalar factor, which in audio processing terms, can be regarded as a change in volume.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2811958,
          "author_name": "max1mum",
          "author_url": "",
          "post_date": "05/14/2024 02:58:52",
          "content": "<p>Thank you very much I was having trouble understanding this one key point! 😁</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2811850": "This is my first BIRDCLEF competition, however, I do not understand why image data performs better than a simple 1d convolution time-series model. Theoretically in image convolution the quality of the notes will be lost when downscaling 10/12/16 bit ADC volume values to 256 or 8 bit colors is this true or am I missing something? Many thanks for answering my question!",
    "2811956": "I am a speech signal processing engineer, and from my point,\n1d time-series signal indeed have much information, but compared with 2d signals, 1d signal is too abstract, thus, for human, directly analyse with it is very challenging. On the other hand, if you have trained a 1d convolution time-series model before, you might find that 1d models are always suffered with overfitting problems and large number of parameters (if with large convolutional kernel ), so if the data number is limited, 2d model may be my first choice.\n\n\"quality of the notes will be lost when downscaling 10/12/16 bit ADC volume values to 256 or 8 bit colors \"\nThere is indeed a loss in quality during downsampling. For instance, if we only consider audio signals with a 16-bit resolution, downsampling or normalizing to 16 bits would involve dividing the audio signal by 2**16, and similarly, downsampling or normalizing to 8 bits would involve dividing the audio signal by 2 **8. Thus, the difference between these two downsampling methods is merely a scalar factor, which in audio processing terms, can be regarded as a change in volume.",
    "2811958": "Thank you very much I was having trouble understanding this one key point! 😁"
  },
  "source": "meta"
}