{
  "id": 179547,
  "title": "Change channel from 3 (rgb) to 1, or optimize the input?",
  "url": "/competitions/birdsong-recognition/discussion/179547",
  "author_name": "",
  "post_date": "2020-09-02T12:58:08.816186700Z",
  "votes": 4,
  "comment_count": 4,
  "views": 0,
  "content": "<p>At the risk of being a stupid question:</p>\n<p>Due to the fact that a spectrogram is not in color/rgb, shouldn't one also change the number of channels to 1, from default 3, before training a model for better performance?<br>\nOr can we use the 3 channels, e.g. add 2 more variants of the spectrogram to the input.<br>\nMono_to_color code seen in public notebooks, seems to be one way to solve it, the best way?</p>",
  "messages": [
    {
      "id": "995444",
      "postDate": "09/02/2020 12:58:08",
      "content": "<p>At the risk of being a stupid question:</p>\n<p>Due to the fact that a spectrogram is not in color/rgb, shouldn't one also change the number of channels to 1, from default 3, before training a model for better performance?<br>\nOr can we use the 3 channels, e.g. add 2 more variants of the spectrogram to the input.<br>\nMono_to_color code seen in public notebooks, seems to be one way to solve it, the best way?</p>",
      "rawMarkdown": "At the risk of being a stupid question:\n\nDue to the fact that a spectrogram is not in color/rgb, shouldn't one also change the number of channels to 1, from default 3, before training a model for better performance?\nOr can we use the 3 channels, e.g. add 2 more variants of the spectrogram to the input.\nMono_to_color code seen in public notebooks, seems to be one way to solve it, the best way?",
      "votes": null
    },
    {
      "id": "995460",
      "postDate": "09/02/2020 13:09:36",
      "content": "<p>I think the real reason why people use 3 channels instead of 1 is the \"Image-net\" or \"noisy-student\" pretrained weights are trained on RGB, and it's common belief that using pretrained weights is better than training from scratch.</p>",
      "rawMarkdown": "I think the real reason why people use 3 channels instead of 1 is the \"Image-net\" or \"noisy-student\" pretrained weights are trained on RGB, and it's common belief that using pretrained weights is better than training from scratch.",
      "votes": null
    },
    {
      "id": "996229",
      "postDate": "09/03/2020 06:49:53",
      "content": "<p>The other good way to make use of 3 channels is add delta and delta delta. I believe it's very common way in audio classification task.</p>",
      "rawMarkdown": "The other good way to make use of 3 channels is add delta and delta delta. I believe it's very common way in audio classification task.",
      "votes": null
    },
    {
      "id": "996273",
      "postDate": "09/03/2020 07:29:56",
      "content": "<p>OK thanks, will look into that!</p>\n<p>Also found an example in a research paper from a university that changed the channel from 3 to 1 to different models.<br>\n\"the number of channels in ResNet has been reduced to 1 (because spectrograms do not have RGB colors); as the number of input features decreased, the model complexity can also be lowered\"<br>\nInception-v3: \"The same weight initialization that<br>\nwas applied to ResNet has also been applied here. The number of channels was<br>\ndecreased before the data flows into Inception module 5c.\"</p>\n<p>Inception-v3 outperformed the others,  due to the number of parameters.<br>\nBut little confused if they only changed the resnet's parameters/weight, Inception parm size look the same but not for resnet.<br>\n\"In our experiment, the parameter size of ResNet-18 is<br>\n9.12 MB, the parameter size of ResNet-34 is 10.41 MB, and the parameter size<br>\nof Inception-v3 is 92.3 MB.\"</p>\n<p>In all cases, still interesting to read some testing on the subject.</p>\n<p>Other interesting reading, they also use these Spec. param.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F9b620b9faafb03201f3b8bae8a624dc2%2FSkrmklipp.jpg?generation=1599118018387236&amp;alt=media\" alt=\"\"></p>\n<p>\"Bird Sound Classification using Convolutional<br>\nNeural Networks\"<br>\n<a href=\"http://ceur-ws.org/Vol-2380/paper_68.pdf\" target=\"_blank\">http://ceur-ws.org/Vol-2380/paper_68.pdf</a></p>",
      "rawMarkdown": "OK thanks, will look into that!\n\nAlso found an example in a research paper from a university that changed the channel from 3 to 1 to different models.\n\"the number of channels in ResNet has been reduced to 1 (because spectrograms do not have RGB colors); as the number of input features decreased, the model complexity can also be lowered\"\nInception-v3: \"The same weight initialization that\nwas applied to ResNet has also been applied here. The number of channels was\ndecreased before the data flows into Inception module 5c.\"\n\nInception-v3 outperformed the others,  due to the number of parameters.\nBut little confused if they only changed the resnet's parameters/weight, Inception parm size look the same but not for resnet.\n\"In our experiment, the parameter size of ResNet-18 is\n9.12 MB, the parameter size of ResNet-34 is 10.41 MB, and the parameter size\nof Inception-v3 is 92.3 MB.\"\n\nIn all cases, still interesting to read some testing on the subject.\n\nOther interesting reading, they also use these Spec. param.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F9b620b9faafb03201f3b8bae8a624dc2%2FSkrmklipp.jpg?generation=1599118018387236&alt=media)\n\n\"Bird Sound Classification using Convolutional\nNeural Networks\"\nhttp://ceur-ws.org/Vol-2380/paper_68.pdf",
      "votes": null
    },
    {
      "id": "1000133",
      "postDate": "09/06/2020 10:18:23",
      "content": "<p>What I do in Pytorch is to input a BxHxW (B: batch size, H:Height, W: width) tensor to my model, and I create the 3 channels in the forward method without duplicating data:</p>\n<pre><code>    def forward(self, x):\n        x = x.unsqueeze(1)\n        x = x.expand(-1, 3, -1, -1)\n        ...\n</code></pre>",
      "rawMarkdown": "What I do in Pytorch is to input a BxHxW (B: batch size, H:Height, W: width) tensor to my model, and I create the 3 channels in the forward method without duplicating data:\n\n```\n    def forward(self, x):\n        x = x.unsqueeze(1)\n        x = x.expand(-1, 3, -1, -1)\n        ...\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 995460,
      "author_name": "quandapro",
      "author_url": "",
      "post_date": "09/02/2020 13:09:36",
      "content": "<p>I think the real reason why people use 3 channels instead of 1 is the \"Image-net\" or \"noisy-student\" pretrained weights are trained on RGB, and it's common belief that using pretrained weights is better than training from scratch.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 996229,
      "author_name": "bamps53",
      "author_url": "",
      "post_date": "09/03/2020 06:49:53",
      "content": "<p>The other good way to make use of 3 channels is add delta and delta delta. I believe it's very common way in audio classification task.</p>",
      "votes": null,
      "replies": [
        {
          "id": 996273,
          "author_name": "kirderf",
          "author_url": "",
          "post_date": "09/03/2020 07:29:56",
          "content": "<p>OK thanks, will look into that!</p>\n<p>Also found an example in a research paper from a university that changed the channel from 3 to 1 to different models.<br>\n\"the number of channels in ResNet has been reduced to 1 (because spectrograms do not have RGB colors); as the number of input features decreased, the model complexity can also be lowered\"<br>\nInception-v3: \"The same weight initialization that<br>\nwas applied to ResNet has also been applied here. The number of channels was<br>\ndecreased before the data flows into Inception module 5c.\"</p>\n<p>Inception-v3 outperformed the others,  due to the number of parameters.<br>\nBut little confused if they only changed the resnet's parameters/weight, Inception parm size look the same but not for resnet.<br>\n\"In our experiment, the parameter size of ResNet-18 is<br>\n9.12 MB, the parameter size of ResNet-34 is 10.41 MB, and the parameter size<br>\nof Inception-v3 is 92.3 MB.\"</p>\n<p>In all cases, still interesting to read some testing on the subject.</p>\n<p>Other interesting reading, they also use these Spec. param.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F9b620b9faafb03201f3b8bae8a624dc2%2FSkrmklipp.jpg?generation=1599118018387236&amp;alt=media\" alt=\"\"></p>\n<p>\"Bird Sound Classification using Convolutional<br>\nNeural Networks\"<br>\n<a href=\"http://ceur-ws.org/Vol-2380/paper_68.pdf\" target=\"_blank\">http://ceur-ws.org/Vol-2380/paper_68.pdf</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1000133,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "09/06/2020 10:18:23",
      "content": "<p>What I do in Pytorch is to input a BxHxW (B: batch size, H:Height, W: width) tensor to my model, and I create the 3 channels in the forward method without duplicating data:</p>\n<pre><code>    def forward(self, x):\n        x = x.unsqueeze(1)\n        x = x.expand(-1, 3, -1, -1)\n        ...\n</code></pre>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "995444": "At the risk of being a stupid question:\n\nDue to the fact that a spectrogram is not in color/rgb, shouldn't one also change the number of channels to 1, from default 3, before training a model for better performance?\nOr can we use the 3 channels, e.g. add 2 more variants of the spectrogram to the input.\nMono_to_color code seen in public notebooks, seems to be one way to solve it, the best way?",
    "995460": "I think the real reason why people use 3 channels instead of 1 is the \"Image-net\" or \"noisy-student\" pretrained weights are trained on RGB, and it's common belief that using pretrained weights is better than training from scratch.",
    "996229": "The other good way to make use of 3 channels is add delta and delta delta. I believe it's very common way in audio classification task.",
    "996273": "OK thanks, will look into that!\n\nAlso found an example in a research paper from a university that changed the channel from 3 to 1 to different models.\n\"the number of channels in ResNet has been reduced to 1 (because spectrograms do not have RGB colors); as the number of input features decreased, the model complexity can also be lowered\"\nInception-v3: \"The same weight initialization that\nwas applied to ResNet has also been applied here. The number of channels was\ndecreased before the data flows into Inception module 5c.\"\n\nInception-v3 outperformed the others,  due to the number of parameters.\nBut little confused if they only changed the resnet's parameters/weight, Inception parm size look the same but not for resnet.\n\"In our experiment, the parameter size of ResNet-18 is\n9.12 MB, the parameter size of ResNet-34 is 10.41 MB, and the parameter size\nof Inception-v3 is 92.3 MB.\"\n\nIn all cases, still interesting to read some testing on the subject.\n\nOther interesting reading, they also use these Spec. param.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F9b620b9faafb03201f3b8bae8a624dc2%2FSkrmklipp.jpg?generation=1599118018387236&alt=media)\n\n\"Bird Sound Classification using Convolutional\nNeural Networks\"\nhttp://ceur-ws.org/Vol-2380/paper_68.pdf",
    "1000133": "What I do in Pytorch is to input a BxHxW (B: batch size, H:Height, W: width) tensor to my model, and I create the 3 channels in the forward method without duplicating data:\n\n```\n    def forward(self, x):\n        x = x.unsqueeze(1)\n        x = x.expand(-1, 3, -1, -1)\n        ...\n```"
  },
  "source": "meta"
}