{
  "id": 69069,
  "title": "Handling 4 channel input",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/69069",
  "author_name": "",
  "post_date": "2018-10-20T05:42:04.956020800Z",
  "votes": 14,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I'd like to discuss whats best to handle 4 Channel input. I quickly checked kernels and found two approches:</p>\n\n<ul>\n<li>Having two branches (3 channel / 1 Channel)</li>\n<li>adding a channel to other channels and then take 3 channel input</li>\n</ul>\n\n<p>First thing that came to my mind was to create a Conv Layer that inputs (size,size,4) and outputs (size,size,3). Anyone experimented with this already?</p>",
  "messages": [
    {
      "id": "406985",
      "postDate": "10/20/2018 05:42:04",
      "content": "<p>I'd like to discuss whats best to handle 4 Channel input. I quickly checked kernels and found two approches:</p>\n\n<ul>\n<li>Having two branches (3 channel / 1 Channel)</li>\n<li>adding a channel to other channels and then take 3 channel input</li>\n</ul>\n\n<p>First thing that came to my mind was to create a Conv Layer that inputs (size,size,4) and outputs (size,size,3). Anyone experimented with this already?</p>",
      "rawMarkdown": "I'd like to discuss whats best to handle 4 Channel input. I quickly checked kernels and found two approches:\n\n - Having two branches (3 channel / 1 Channel)\n - adding a channel to other channels and then take 3 channel input\n\nFirst thing that came to my mind was to create a Conv Layer that inputs (size,size,4) and outputs (size,size,3). Anyone experimented with this already?",
      "votes": null
    },
    {
      "id": "407000",
      "postDate": "10/20/2018 06:30:53",
      "content": "<p>I'm working on a way to modify pretrained model to be usable with 4 channels. Keypoints of my approach:</p>\n\n<ul>\n<li>Take the structure of some standard model (e.g. InceptionResnetV2)</li>\n<li>Modify the input layer from (299, 299, 3) to (w, h, 4)</li>\n<li>Take the weights of a pretrained model (all layers will have same shape of parameters except the first one)</li>\n<li>Modify the first layer weights to the desired shape (You need to calculate 4th channel weights somehow: I'm currently experimenting between random initialization, mean of the 3 channels parameters and taking same as one channel)</li>\n</ul>\n\n<p>I have some success with this approach (cv 0.364 after 2 epochs, didn't try on LB yet), but the training is slower.</p>\n\n<p>This is btw also a way how to use full size input (512, 512, 4) instead (299, 299, 3).</p>",
      "rawMarkdown": "I'm working on a way to modify pretrained model to be usable with 4 channels. Keypoints of my approach:\n\n - Take the structure of some standard model (e.g. InceptionResnetV2)\n - Modify the input layer from (299, 299, 3) to (w, h, 4)\n - Take the weights of a pretrained model (all layers will have same shape of parameters except the first one)\n - Modify the first layer weights to the desired shape (You need to calculate 4th channel weights somehow: I'm currently experimenting between random initialization, mean of the 3 channels parameters and taking same as one channel)\n\nI have some success with this approach (cv 0.364 after 2 epochs, didn't try on LB yet), but the training is slower.\n\nThis is btw also a way how to use full size input (512, 512, 4) instead (299, 299, 3).",
      "votes": null
    },
    {
      "id": "407017",
      "postDate": "10/20/2018 07:43:11",
      "content": "<p>What about multiplying green channel with each of red, blue , yellow?</p>",
      "rawMarkdown": "What about multiplying green channel with each of red, blue , yellow?",
      "votes": null
    },
    {
      "id": "407214",
      "postDate": "10/20/2018 16:18:14",
      "content": "<p>This was the first thing I tried. A 1x1 convolution that drops from 4 dimensions down to 3 dimensions. I believe this is likely a good way to do it but I am unsure of the best way to train that first layer. Typically people will train the last few layers and freeze the rest to initialize for transfer learning, but I have never seen someone do that for the first layer as well. Seems like it would not be able to pass the gradient all that well. Haven't fully formulated the idea but there is possibly some way you could do a sort of skip connection. </p>\n\n<p>Anyway in my experiments it did ok but not great. Built off one of the kernels and it did worse than that did with just the 3 channels. </p>",
      "rawMarkdown": "This was the first thing I tried. A 1x1 convolution that drops from 4 dimensions down to 3 dimensions. I believe this is likely a good way to do it but I am unsure of the best way to train that first layer. Typically people will train the last few layers and freeze the rest to initialize for transfer learning, but I have never seen someone do that for the first layer as well. Seems like it would not be able to pass the gradient all that well. Haven't fully formulated the idea but there is possibly some way you could do a sort of skip connection. \n\nAnyway in my experiments it did ok but not great. Built off one of the kernels and it did worse than that did with just the 3 channels.",
      "votes": null
    },
    {
      "id": "407248",
      "postDate": "10/20/2018 17:50:28",
      "content": "<p>Initially I started with 4 channels, RGBY. After much research all the previous examples I could find for this sort of task used either RG or RGB. My current model I am discarding Y and using only RGB. It runs faster and the result seems to be the same.</p>",
      "rawMarkdown": "Initially I started with 4 channels, RGBY. After much research all the previous examples I could find for this sort of task used either RG or RGB. My current model I am discarding Y and using only RGB. It runs faster and the result seems to be the same.",
      "votes": null
    },
    {
      "id": "407250",
      "postDate": "10/20/2018 17:55:36",
      "content": "<p>I've applied something similar to work on a model that was intended for 64x64x?. I resized the inputs to 256x256x3 (discard yellow), and added a convolution layer to match the input. The convolution layer I used has a 64x64 kernel and 256 filters. So far it is performing well, training without pretrained weights.</p>",
      "rawMarkdown": "I've applied something similar to work on a model that was intended for 64x64x?. I resized the inputs to 256x256x3 (discard yellow), and added a convolution layer to match the input. The convolution layer I used has a 64x64 kernel and 256 filters. So far it is performing well, training without pretrained weights.",
      "votes": null
    },
    {
      "id": "407346",
      "postDate": "10/20/2018 23:57:22",
      "content": "<p>You can replace the first layer 3 channel convolution by 4 channel one and initialize additional weights for the 4-th channel by zero. In this case the initial state of the network is the same as the pretrained one; however, the model now is capable to incorporation features from Y channel into prediction. You can check how it is done in this kernel: <a href=\"https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-fast-ai\">https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-fast-ai</a> </p>",
      "rawMarkdown": "You can replace the first layer 3 channel convolution by 4 channel one and initialize additional weights for the 4-th channel by zero. In this case the initial state of the network is the same as the pretrained one; however, the model now is capable to incorporation features from Y channel into prediction. You can check how it is done in this kernel: https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-fast-ai",
      "votes": null
    },
    {
      "id": "407430",
      "postDate": "10/21/2018 06:24:12",
      "content": "<p>I added simple ConvBlock reducing from 4 to 3 channels before pretrained model. Works good.\n<a href=\"/ryches\">@ryches</a> I freeze pretrained network, and set this Convlayer and last layers to trainable</p>",
      "rawMarkdown": "I added simple ConvBlock reducing from 4 to 3 channels before pretrained model. Works good.\n@ryches I freeze pretrained network, and set this Convlayer and last layers to trainable",
      "votes": null
    },
    {
      "id": "407908",
      "postDate": "10/22/2018 01:24:42",
      "content": "<p>Hi, Dieter. Did you freeze the pretained network throughout the training or just the a few epochs from the beginning.</p>",
      "rawMarkdown": "Hi, Dieter. Did you freeze the pretained network throughout the training or just the a few epochs from the beginning.",
      "votes": null
    },
    {
      "id": "409585",
      "postDate": "10/24/2018 14:32:41",
      "content": "<p>Comment in one of the discussions from the host that yellow is seldom used.  I took this to mean there may be one or more of the rare proteins where yellow will be helpful.  Like you discarding yellow for now, but watching for size of confusion error on the proteins with limited number of images.</p>",
      "rawMarkdown": "Comment in one of the discussions from the host that yellow is seldom used.  I took this to mean there may be one or more of the rare proteins where yellow will be helpful.  Like you discarding yellow for now, but watching for size of confusion error on the proteins with limited number of images.",
      "votes": null
    },
    {
      "id": "409590",
      "postDate": "10/24/2018 14:42:08",
      "content": "<p>When using lights (computer monitors included) yellow is created by mixing red and green.  Would make sense to do additive of yellow into the red and green channels.</p>\n\n<p>When using paints yellow is actually the minus-blue primary - red and green paint do not produce yellow.  So I guess here a subtraction from blue.</p>\n\n<p>When doing the protein process my sense is that the four colors are the result of different staining and lighting - so not sure that either of the above make sense.  </p>",
      "rawMarkdown": "When using lights (computer monitors included) yellow is created by mixing red and green.  Would make sense to do additive of yellow into the red and green channels.\n\nWhen using paints yellow is actually the minus-blue primary - red and green paint do not produce yellow.  So I guess here a subtraction from blue.\n\nWhen doing the protein process my sense is that the four colors are the result of different staining and lighting - so not sure that either of the above make sense.",
      "votes": null
    },
    {
      "id": "409644",
      "postDate": "10/24/2018 16:28:58",
      "content": "<p>I don't think we can just combine these 'color' channels the way you combine paints or lights.  The colors assigned could be completely arbitrary, and even if they were not there is no guarantee that they would combine in the way that you combine paints.</p>",
      "rawMarkdown": "I don't think we can just combine these 'color' channels the way you combine paints or lights.  The colors assigned could be completely arbitrary, and even if they were not there is no guarantee that they would combine in the way that you combine paints.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 407000,
      "author_name": "rejpalcz",
      "author_url": "",
      "post_date": "10/20/2018 06:30:53",
      "content": "<p>I'm working on a way to modify pretrained model to be usable with 4 channels. Keypoints of my approach:</p>\n\n<ul>\n<li>Take the structure of some standard model (e.g. InceptionResnetV2)</li>\n<li>Modify the input layer from (299, 299, 3) to (w, h, 4)</li>\n<li>Take the weights of a pretrained model (all layers will have same shape of parameters except the first one)</li>\n<li>Modify the first layer weights to the desired shape (You need to calculate 4th channel weights somehow: I'm currently experimenting between random initialization, mean of the 3 channels parameters and taking same as one channel)</li>\n</ul>\n\n<p>I have some success with this approach (cv 0.364 after 2 epochs, didn't try on LB yet), but the training is slower.</p>\n\n<p>This is btw also a way how to use full size input (512, 512, 4) instead (299, 299, 3).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 407017,
      "author_name": "christofhenkel",
      "author_url": "",
      "post_date": "10/20/2018 07:43:11",
      "content": "<p>What about multiplying green channel with each of red, blue , yellow?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 407214,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "10/20/2018 16:18:14",
      "content": "<p>This was the first thing I tried. A 1x1 convolution that drops from 4 dimensions down to 3 dimensions. I believe this is likely a good way to do it but I am unsure of the best way to train that first layer. Typically people will train the last few layers and freeze the rest to initialize for transfer learning, but I have never seen someone do that for the first layer as well. Seems like it would not be able to pass the gradient all that well. Haven't fully formulated the idea but there is possibly some way you could do a sort of skip connection. </p>\n\n<p>Anyway in my experiments it did ok but not great. Built off one of the kernels and it did worse than that did with just the 3 channels. </p>",
      "votes": null,
      "replies": [
        {
          "id": 407250,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "10/20/2018 17:55:36",
          "content": "<p>I've applied something similar to work on a model that was intended for 64x64x?. I resized the inputs to 256x256x3 (discard yellow), and added a convolution layer to match the input. The convolution layer I used has a 64x64 kernel and 256 filters. So far it is performing well, training without pretrained weights.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 407248,
      "author_name": "ldm314",
      "author_url": "",
      "post_date": "10/20/2018 17:50:28",
      "content": "<p>Initially I started with 4 channels, RGBY. After much research all the previous examples I could find for this sort of task used either RG or RGB. My current model I am discarding Y and using only RGB. It runs faster and the result seems to be the same.</p>",
      "votes": null,
      "replies": [
        {
          "id": 409585,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "10/24/2018 14:32:41",
          "content": "<p>Comment in one of the discussions from the host that yellow is seldom used.  I took this to mean there may be one or more of the rare proteins where yellow will be helpful.  Like you discarding yellow for now, but watching for size of confusion error on the proteins with limited number of images.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 407346,
      "author_name": "iafoss",
      "author_url": "",
      "post_date": "10/20/2018 23:57:22",
      "content": "<p>You can replace the first layer 3 channel convolution by 4 channel one and initialize additional weights for the 4-th channel by zero. In this case the initial state of the network is the same as the pretrained one; however, the model now is capable to incorporation features from Y channel into prediction. You can check how it is done in this kernel: <a href=\"https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-fast-ai\">https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-fast-ai</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 407430,
      "author_name": "christofhenkel",
      "author_url": "",
      "post_date": "10/21/2018 06:24:12",
      "content": "<p>I added simple ConvBlock reducing from 4 to 3 channels before pretrained model. Works good.\n<a href=\"/ryches\">@ryches</a> I freeze pretrained network, and set this Convlayer and last layers to trainable</p>",
      "votes": null,
      "replies": [
        {
          "id": 407908,
          "author_name": "garybios",
          "author_url": "",
          "post_date": "10/22/2018 01:24:42",
          "content": "<p>Hi, Dieter. Did you freeze the pretained network throughout the training or just the a few epochs from the beginning.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 409590,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "10/24/2018 14:42:08",
      "content": "<p>When using lights (computer monitors included) yellow is created by mixing red and green.  Would make sense to do additive of yellow into the red and green channels.</p>\n\n<p>When using paints yellow is actually the minus-blue primary - red and green paint do not produce yellow.  So I guess here a subtraction from blue.</p>\n\n<p>When doing the protein process my sense is that the four colors are the result of different staining and lighting - so not sure that either of the above make sense.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 409644,
          "author_name": "robertkag",
          "author_url": "",
          "post_date": "10/24/2018 16:28:58",
          "content": "<p>I don't think we can just combine these 'color' channels the way you combine paints or lights.  The colors assigned could be completely arbitrary, and even if they were not there is no guarantee that they would combine in the way that you combine paints.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "406985": "I'd like to discuss whats best to handle 4 Channel input. I quickly checked kernels and found two approches:\n\n - Having two branches (3 channel / 1 Channel)\n - adding a channel to other channels and then take 3 channel input\n\nFirst thing that came to my mind was to create a Conv Layer that inputs (size,size,4) and outputs (size,size,3). Anyone experimented with this already?",
    "407000": "I'm working on a way to modify pretrained model to be usable with 4 channels. Keypoints of my approach:\n\n - Take the structure of some standard model (e.g. InceptionResnetV2)\n - Modify the input layer from (299, 299, 3) to (w, h, 4)\n - Take the weights of a pretrained model (all layers will have same shape of parameters except the first one)\n - Modify the first layer weights to the desired shape (You need to calculate 4th channel weights somehow: I'm currently experimenting between random initialization, mean of the 3 channels parameters and taking same as one channel)\n\nI have some success with this approach (cv 0.364 after 2 epochs, didn't try on LB yet), but the training is slower.\n\nThis is btw also a way how to use full size input (512, 512, 4) instead (299, 299, 3).",
    "407017": "What about multiplying green channel with each of red, blue , yellow?",
    "407214": "This was the first thing I tried. A 1x1 convolution that drops from 4 dimensions down to 3 dimensions. I believe this is likely a good way to do it but I am unsure of the best way to train that first layer. Typically people will train the last few layers and freeze the rest to initialize for transfer learning, but I have never seen someone do that for the first layer as well. Seems like it would not be able to pass the gradient all that well. Haven't fully formulated the idea but there is possibly some way you could do a sort of skip connection. \n\nAnyway in my experiments it did ok but not great. Built off one of the kernels and it did worse than that did with just the 3 channels.",
    "407248": "Initially I started with 4 channels, RGBY. After much research all the previous examples I could find for this sort of task used either RG or RGB. My current model I am discarding Y and using only RGB. It runs faster and the result seems to be the same.",
    "407250": "I've applied something similar to work on a model that was intended for 64x64x?. I resized the inputs to 256x256x3 (discard yellow), and added a convolution layer to match the input. The convolution layer I used has a 64x64 kernel and 256 filters. So far it is performing well, training without pretrained weights.",
    "407346": "You can replace the first layer 3 channel convolution by 4 channel one and initialize additional weights for the 4-th channel by zero. In this case the initial state of the network is the same as the pretrained one; however, the model now is capable to incorporation features from Y channel into prediction. You can check how it is done in this kernel: https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-fast-ai",
    "407430": "I added simple ConvBlock reducing from 4 to 3 channels before pretrained model. Works good.\n@ryches I freeze pretrained network, and set this Convlayer and last layers to trainable",
    "407908": "Hi, Dieter. Did you freeze the pretained network throughout the training or just the a few epochs from the beginning.",
    "409585": "Comment in one of the discussions from the host that yellow is seldom used.  I took this to mean there may be one or more of the rare proteins where yellow will be helpful.  Like you discarding yellow for now, but watching for size of confusion error on the proteins with limited number of images.",
    "409590": "When using lights (computer monitors included) yellow is created by mixing red and green.  Would make sense to do additive of yellow into the red and green channels.\n\nWhen using paints yellow is actually the minus-blue primary - red and green paint do not produce yellow.  So I guess here a subtraction from blue.\n\nWhen doing the protein process my sense is that the four colors are the result of different staining and lighting - so not sure that either of the above make sense.",
    "409644": "I don't think we can just combine these 'color' channels the way you combine paints or lights.  The colors assigned could be completely arbitrary, and even if they were not there is no guarantee that they would combine in the way that you combine paints."
  },
  "source": "meta"
}