{
  "id": 496943,
  "title": "Questions about the input dimensions of the CNN for BirdCLEF2024",
  "url": "/competitions/birdclef-2024/discussion/496943",
  "author_name": "",
  "post_date": "2024-04-23T03:09:58.671973100Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I am a beginner and have a rudimentary question.<br>\nPlease tell me about the input dimension of CNN using pytorch.<br>\nI am planning to use a trained model (resnet18) to evaluate a spectrogram I have created.</p>\n<p>---About spectrograms: -------------------------------------<br>\nI have transformed BirdCLEF audio data into a spectrogram.<br>\nThe spectrogram is an image whose channels are one-dimensional.<br>\nInput dimension (batch, ch, H, W) = ( *, 1, *, *)</p>\n<h2>\"*\" is an arbitrary value.  </h2>\n<p>Only the output dimension is changed from class 1000 to the BirdCLEF2024 class label,<br>\nI thought to train with \"requires_grad=False\" for all but the output layer,<br>\nHowever, since the spectrogram channel is 1-dimensional, the input channel of resnet18 also needs to be changed from 3 to 1, and the input channel also needs to be set \"requires_grad=True\".</p>\n<p><strong>So, my question is.</strong></p>\n<ol>\n<li>If I expand the input dimension (batch, ch, H, W) = ( *, 1, *, *) to expand(-1, 3, -1, -1, -1), will it not worsen the accuracy of the model training?</li>\n<li>When using the input dimension as one channel, is it necessary to set \"requires_grad = True\" for the input and output layers, or is it better to set \"requires_grad = True\" for all layers for fine tuning?</li>\n</ol>\n<p>I would appreciate it if you could tell me.</p>",
  "messages": [
    {
      "id": "2768779",
      "postDate": "04/23/2024 03:09:58",
      "content": "<p>I am a beginner and have a rudimentary question.<br>\nPlease tell me about the input dimension of CNN using pytorch.<br>\nI am planning to use a trained model (resnet18) to evaluate a spectrogram I have created.</p>\n<p>---About spectrograms: -------------------------------------<br>\nI have transformed BirdCLEF audio data into a spectrogram.<br>\nThe spectrogram is an image whose channels are one-dimensional.<br>\nInput dimension (batch, ch, H, W) = ( *, 1, *, *)</p>\n<h2>\"*\" is an arbitrary value.  </h2>\n<p>Only the output dimension is changed from class 1000 to the BirdCLEF2024 class label,<br>\nI thought to train with \"requires_grad=False\" for all but the output layer,<br>\nHowever, since the spectrogram channel is 1-dimensional, the input channel of resnet18 also needs to be changed from 3 to 1, and the input channel also needs to be set \"requires_grad=True\".</p>\n<p><strong>So, my question is.</strong></p>\n<ol>\n<li>If I expand the input dimension (batch, ch, H, W) = ( *, 1, *, *) to expand(-1, 3, -1, -1, -1), will it not worsen the accuracy of the model training?</li>\n<li>When using the input dimension as one channel, is it necessary to set \"requires_grad = True\" for the input and output layers, or is it better to set \"requires_grad = True\" for all layers for fine tuning?</li>\n</ol>\n<p>I would appreciate it if you could tell me.</p>",
      "rawMarkdown": "I am a beginner and have a rudimentary question.\nPlease tell me about the input dimension of CNN using pytorch.\nI am planning to use a trained model (resnet18) to evaluate a spectrogram I have created.\n\n---About spectrograms: -------------------------------------\nI have transformed BirdCLEF audio data into a spectrogram.\nThe spectrogram is an image whose channels are one-dimensional.\nInput dimension (batch, ch, H, W) = ( *, 1, *, *)\n\"*\" is an arbitrary value.  \n---------------------------------------------------------------------\n\nOnly the output dimension is changed from class 1000 to the BirdCLEF2024 class label,\nI thought to train with \"requires_grad=False\" for all but the output layer,\nHowever, since the spectrogram channel is 1-dimensional, the input channel of resnet18 also needs to be changed from 3 to 1, and the input channel also needs to be set \"requires_grad=True\".\n\n**So, my question is.**\n1. If I expand the input dimension (batch, ch, H, W) = ( *, 1, *, *) to expand(-1, 3, -1, -1, -1), will it not worsen the accuracy of the model training?\n2. When using the input dimension as one channel, is it necessary to set \"requires_grad = True\" for the input and output layers, or is it better to set \"requires_grad = True\" for all layers for fine tuning?\n\nI would appreciate it if you could tell me.",
      "votes": null
    },
    {
      "id": "2769423",
      "postDate": "04/23/2024 10:42:32",
      "content": "<p>Your pretrained model requires 3 channels, therefore you need to provide 3 channel images as input. <code>expand(-1, 3, -1, -1, -1)</code> is a good way to do it.</p>\n<p>For freezing layers or not, there is only one way to answer your question: try both and see what works best.</p>",
      "rawMarkdown": "Your pretrained model requires 3 channels, therefore you need to provide 3 channel images as input. ` expand(-1, 3, -1, -1, -1)` is a good way to do it.\n\nFor freezing layers or not, there is only one way to answer your question: try both and see what works best.",
      "votes": null
    },
    {
      "id": "2769544",
      "postDate": "04/23/2024 12:10:28",
      "content": "<p>Thank you for telling us about it.<br>\nI try to model both method using cross-validation as you suggested.</p>",
      "rawMarkdown": "Thank you for telling us about it.\nI try to model both method using cross-validation as you suggested.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2769423,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "04/23/2024 10:42:32",
      "content": "<p>Your pretrained model requires 3 channels, therefore you need to provide 3 channel images as input. <code>expand(-1, 3, -1, -1, -1)</code> is a good way to do it.</p>\n<p>For freezing layers or not, there is only one way to answer your question: try both and see what works best.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2769544,
          "author_name": "madao13",
          "author_url": "",
          "post_date": "04/23/2024 12:10:28",
          "content": "<p>Thank you for telling us about it.<br>\nI try to model both method using cross-validation as you suggested.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2768779": "I am a beginner and have a rudimentary question.\nPlease tell me about the input dimension of CNN using pytorch.\nI am planning to use a trained model (resnet18) to evaluate a spectrogram I have created.\n\n---About spectrograms: -------------------------------------\nI have transformed BirdCLEF audio data into a spectrogram.\nThe spectrogram is an image whose channels are one-dimensional.\nInput dimension (batch, ch, H, W) = ( *, 1, *, *)\n\"*\" is an arbitrary value.  \n---------------------------------------------------------------------\n\nOnly the output dimension is changed from class 1000 to the BirdCLEF2024 class label,\nI thought to train with \"requires_grad=False\" for all but the output layer,\nHowever, since the spectrogram channel is 1-dimensional, the input channel of resnet18 also needs to be changed from 3 to 1, and the input channel also needs to be set \"requires_grad=True\".\n\n**So, my question is.**\n1. If I expand the input dimension (batch, ch, H, W) = ( *, 1, *, *) to expand(-1, 3, -1, -1, -1), will it not worsen the accuracy of the model training?\n2. When using the input dimension as one channel, is it necessary to set \"requires_grad = True\" for the input and output layers, or is it better to set \"requires_grad = True\" for all layers for fine tuning?\n\nI would appreciate it if you could tell me.",
    "2769423": "Your pretrained model requires 3 channels, therefore you need to provide 3 channel images as input. ` expand(-1, 3, -1, -1, -1)` is a good way to do it.\n\nFor freezing layers or not, there is only one way to answer your question: try both and see what works best.",
    "2769544": "Thank you for telling us about it.\nI try to model both method using cross-validation as you suggested."
  },
  "source": "meta"
}