{
  "id": 202065,
  "title": "Data Augmentation Question - Channel Wise or Globally",
  "url": "/competitions/rfcx-species-audio-detection/discussion/202065",
  "author_name": "",
  "post_date": "2020-12-08T06:09:46.377220300Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>When using models like ResNeSt50 we stack the input spectrogram 3 times, such that we get an image with 3 channels.</p>\n<p>Do we apply data augmentation to each spectrogram in each respective channel? or do we apply data augmentation such that the spectrograms in each channel have the exact same augmentation?</p>\n<p>So for example, if I apply gaussian noise, I could apply it first to the spectrogram and then stack that spectrogram three times, or I could first stack and then apply gaussian noise to each one of them, such that each one has a different noise.</p>\n<p>In the latter case, would it be damaging or beneficial to the learned features? That's primarily my concern, I did some experimenting but it doesn't seem like it matters much, but then again again there might be some paper about this.</p>\n<p>I hope it's clear what I'm asking?</p>",
  "messages": [
    {
      "id": "1105715",
      "postDate": "12/08/2020 06:09:46",
      "content": "<p>When using models like ResNeSt50 we stack the input spectrogram 3 times, such that we get an image with 3 channels.</p>\n<p>Do we apply data augmentation to each spectrogram in each respective channel? or do we apply data augmentation such that the spectrograms in each channel have the exact same augmentation?</p>\n<p>So for example, if I apply gaussian noise, I could apply it first to the spectrogram and then stack that spectrogram three times, or I could first stack and then apply gaussian noise to each one of them, such that each one has a different noise.</p>\n<p>In the latter case, would it be damaging or beneficial to the learned features? That's primarily my concern, I did some experimenting but it doesn't seem like it matters much, but then again again there might be some paper about this.</p>\n<p>I hope it's clear what I'm asking?</p>",
      "rawMarkdown": "When using models like ResNeSt50 we stack the input spectrogram 3 times, such that we get an image with 3 channels.\n\nDo we apply data augmentation to each spectrogram in each respective channel? or do we apply data augmentation such that the spectrograms in each channel have the exact same augmentation?\n\nSo for example, if I apply gaussian noise, I could apply it first to the spectrogram and then stack that spectrogram three times, or I could first stack and then apply gaussian noise to each one of them, such that each one has a different noise.\n\nIn the latter case, would it be damaging or beneficial to the learned features? That's primarily my concern, I did some experimenting but it doesn't seem like it matters much, but then again again there might be some paper about this.\n\nI hope it's clear what I'm asking?",
      "votes": null
    },
    {
      "id": "1106253",
      "postDate": "12/08/2020 16:49:35",
      "content": "<p>In ResNet you only need one image, however stacking is a good hack if you don't know how to use 1 channel or how to colorize. In your case, guassian noise makes some sense if the image is crisper because image classifiers are more based on texture than other things. However, there is a more creative way to this by doing a 2D convolution that outputs to a new array with a new channel of 3.   </p>",
      "rawMarkdown": "In ResNet you only need one image, however stacking is a good hack if you don't know how to use 1 channel or how to colorize. In your case, guassian noise makes some sense if the image is crisper because image classifiers are more based on texture than other things. However, there is a more creative way to this by doing a 2D convolution that outputs to a new array with a new channel of 3.",
      "votes": null
    },
    {
      "id": "1106594",
      "postDate": "12/09/2020 00:09:54",
      "content": "<p>I could be wrong, but here's my take. </p>\n<p>I think the better method is applying the same noise to each channel. In reality, there's only 1 spectrogram: We just stack it to match the required ResNet input. By changing the noise for each of the 3 channels, the model will try to learn the pattern for why there are difference between the 3 channels, but the only differences will be the artificial noise you added to the data.</p>\n<p>As you pointed out, there's not much of a difference between the two approaches, but in theory, I think applying the same noise to each of the 3 channels is better.</p>",
      "rawMarkdown": "I could be wrong, but here's my take. \n\nI think the better method is applying the same noise to each channel. In reality, there's only 1 spectrogram: We just stack it to match the required ResNet input. By changing the noise for each of the 3 channels, the model will try to learn the pattern for why there are difference between the 3 channels, but the only differences will be the artificial noise you added to the data.\n\nAs you pointed out, there's not much of a difference between the two approaches, but in theory, I think applying the same noise to each of the 3 channels is better.",
      "votes": null
    },
    {
      "id": "1110409",
      "postDate": "12/12/2020 18:05:15",
      "content": "<p>FWIW, we can modify the first input layer of the prepackages models so that it accepts in channel inputs.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1557202%2F1ff20aa1d06d2c8cee6537beddb2a270%2FScreenshot%20from%202020-12-12%2011-59-54.png?generation=1607796045404271&amp;alt=media\" alt=\"\"></p>\n<p>Although duplicating 1 channel data to 3 channel data does some redundant computation, I think there is really no difference in model behavior compared to the above approach, since 3 weights dot product with 3 identical numbers in 3 channels does loosely speaking the same thing as multiplying 1  weight to 1 number in one channel.</p>",
      "rawMarkdown": "FWIW, we can modify the first input layer of the prepackages models so that it accepts in channel inputs.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1557202%2F1ff20aa1d06d2c8cee6537beddb2a270%2FScreenshot%20from%202020-12-12%2011-59-54.png?generation=1607796045404271&alt=media)\n\nAlthough duplicating 1 channel data to 3 channel data does some redundant computation, I think there is really no difference in model behavior compared to the above approach, since 3 weights dot product with 3 identical numbers in 3 channels does loosely speaking the same thing as multiplying 1  weight to 1 number in one channel.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1106253,
      "author_name": "eladwar",
      "author_url": "",
      "post_date": "12/08/2020 16:49:35",
      "content": "<p>In ResNet you only need one image, however stacking is a good hack if you don't know how to use 1 channel or how to colorize. In your case, guassian noise makes some sense if the image is crisper because image classifiers are more based on texture than other things. However, there is a more creative way to this by doing a 2D convolution that outputs to a new array with a new channel of 3.   </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1106594,
      "author_name": "maltonji",
      "author_url": "",
      "post_date": "12/09/2020 00:09:54",
      "content": "<p>I could be wrong, but here's my take. </p>\n<p>I think the better method is applying the same noise to each channel. In reality, there's only 1 spectrogram: We just stack it to match the required ResNet input. By changing the noise for each of the 3 channels, the model will try to learn the pattern for why there are difference between the 3 channels, but the only differences will be the artificial noise you added to the data.</p>\n<p>As you pointed out, there's not much of a difference between the two approaches, but in theory, I think applying the same noise to each of the 3 channels is better.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1110409,
      "author_name": "barnwellguy",
      "author_url": "",
      "post_date": "12/12/2020 18:05:15",
      "content": "<p>FWIW, we can modify the first input layer of the prepackages models so that it accepts in channel inputs.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1557202%2F1ff20aa1d06d2c8cee6537beddb2a270%2FScreenshot%20from%202020-12-12%2011-59-54.png?generation=1607796045404271&amp;alt=media\" alt=\"\"></p>\n<p>Although duplicating 1 channel data to 3 channel data does some redundant computation, I think there is really no difference in model behavior compared to the above approach, since 3 weights dot product with 3 identical numbers in 3 channels does loosely speaking the same thing as multiplying 1  weight to 1 number in one channel.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1105715": "When using models like ResNeSt50 we stack the input spectrogram 3 times, such that we get an image with 3 channels.\n\nDo we apply data augmentation to each spectrogram in each respective channel? or do we apply data augmentation such that the spectrograms in each channel have the exact same augmentation?\n\nSo for example, if I apply gaussian noise, I could apply it first to the spectrogram and then stack that spectrogram three times, or I could first stack and then apply gaussian noise to each one of them, such that each one has a different noise.\n\nIn the latter case, would it be damaging or beneficial to the learned features? That's primarily my concern, I did some experimenting but it doesn't seem like it matters much, but then again again there might be some paper about this.\n\nI hope it's clear what I'm asking?",
    "1106253": "In ResNet you only need one image, however stacking is a good hack if you don't know how to use 1 channel or how to colorize. In your case, guassian noise makes some sense if the image is crisper because image classifiers are more based on texture than other things. However, there is a more creative way to this by doing a 2D convolution that outputs to a new array with a new channel of 3.",
    "1106594": "I could be wrong, but here's my take. \n\nI think the better method is applying the same noise to each channel. In reality, there's only 1 spectrogram: We just stack it to match the required ResNet input. By changing the noise for each of the 3 channels, the model will try to learn the pattern for why there are difference between the 3 channels, but the only differences will be the artificial noise you added to the data.\n\nAs you pointed out, there's not much of a difference between the two approaches, but in theory, I think applying the same noise to each of the 3 channels is better.",
    "1110409": "FWIW, we can modify the first input layer of the prepackages models so that it accepts in channel inputs.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1557202%2F1ff20aa1d06d2c8cee6537beddb2a270%2FScreenshot%20from%202020-12-12%2011-59-54.png?generation=1607796045404271&alt=media)\n\nAlthough duplicating 1 channel data to 3 channel data does some redundant computation, I think there is really no difference in model behavior compared to the above approach, since 3 weights dot product with 3 identical numbers in 3 channels does loosely speaking the same thing as multiplying 1  weight to 1 number in one channel."
  },
  "source": "meta"
}