{
  "id": 106715,
  "title": "Best way to use both sites?",
  "url": "/competitions/recursion-cellular-image-classification/discussion/106715",
  "author_name": "",
  "post_date": "2019-08-30T18:11:57.182111600Z",
  "votes": 9,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I tried three different approaches to using the images from site1 and site2:\n* combine them into 512x512 image with 12 channels\n* combine into 1024x512 image with 6 channels\n* use them separately as 512x512 images with 6 channels and apply model twice. Like\n<code>\ndef forward(x):\n  return self.model(x[0]) + self.model(x[1])\n</code></p>\n\n<p>Without changing the model (Resnet18), using 12 channels resulted in worse accuracy. Presumably because only the first conv sees the extra data.</p>\n\n<p>The later two perform about the same, but separating them seems preferable to me, since I can use other model architectures and share weights more easily between the two sites.</p>\n\n<p>Anyone else found a good way to combine the images from the two sites?</p>",
  "messages": [
    {
      "id": "613675",
      "postDate": "08/30/2019 18:11:57",
      "content": "<p>I tried three different approaches to using the images from site1 and site2:\n* combine them into 512x512 image with 12 channels\n* combine into 1024x512 image with 6 channels\n* use them separately as 512x512 images with 6 channels and apply model twice. Like\n<code>\ndef forward(x):\n  return self.model(x[0]) + self.model(x[1])\n</code></p>\n\n<p>Without changing the model (Resnet18), using 12 channels resulted in worse accuracy. Presumably because only the first conv sees the extra data.</p>\n\n<p>The later two perform about the same, but separating them seems preferable to me, since I can use other model architectures and share weights more easily between the two sites.</p>\n\n<p>Anyone else found a good way to combine the images from the two sites?</p>",
      "rawMarkdown": "I tried three different approaches to using the images from site1 and site2:\n* combine them into 512x512 image with 12 channels\n* combine into 1024x512 image with 6 channels\n* use them separately as 512x512 images with 6 channels and apply model twice. Like\n```\ndef forward(x):\n  return self.model(x[0]) + self.model(x[1])\n```\n\nWithout changing the model (Resnet18), using 12 channels resulted in worse accuracy. Presumably because only the first conv sees the extra data.\n\nThe later two perform about the same, but separating them seems preferable to me, since I can use other model architectures and share weights more easily between the two sites.\n\nAnyone else found a good way to combine the images from the two sites?",
      "votes": null
    },
    {
      "id": "613721",
      "postDate": "08/30/2019 18:59:29",
      "content": "<p>I've been treating them as straight up separate images, so the training set becomes ~70K images.</p>",
      "rawMarkdown": "I've been treating them as straight up separate images, so the training set becomes ~70K images.",
      "votes": null
    },
    {
      "id": "615513",
      "postDate": "09/02/2019 03:22:35",
      "content": "<p>Interesting, are you however using both sites to make your predictions on each instance of the test set <a href=\"/interneuron\">@interneuron</a> ?</p>",
      "rawMarkdown": "Interesting, are you however using both sites to make your predictions on each instance of the test set @interneuron ?",
      "votes": null
    },
    {
      "id": "615542",
      "postDate": "09/02/2019 04:33:22",
      "content": "<p>yes, so total test size is ~39K</p>",
      "rawMarkdown": "yes, so total test size is ~39K",
      "votes": null
    },
    {
      "id": "617690",
      "postDate": "09/04/2019 11:48:00",
      "content": "<p>I'm using the 2 sites as separate images. Interestingly, taking the minimum of the 2 predicted probabilities gave me a slightly better LB score than taking the mean. I hadn't even thought of combining them. Maybe something to try, though the image discontinuity at the join may introduce noise that will confuse the model, so I would prefer to leave it as [512, 512, 12] if the two are combined.</p>",
      "rawMarkdown": "I'm using the 2 sites as separate images. Interestingly, taking the minimum of the 2 predicted probabilities gave me a slightly better LB score than taking the mean. I hadn't even thought of combining them. Maybe something to try, though the image discontinuity at the join may introduce noise that will confuse the model, so I would prefer to leave it as [512, 512, 12] if the two are combined.",
      "votes": null
    },
    {
      "id": "617969",
      "postDate": "09/04/2019 17:08:34",
      "content": "<p>I was also thinking about the same. However, I am not sure if just putting the two 6D site images to 12D images is the best option, as the data is not spatial linked for the two images, i.e., is from different spots on the well plate.</p>",
      "rawMarkdown": "I was also thinking about the same. However, I am not sure if just putting the two 6D site images to 12D images is the best option, as the data is not spatial linked for the two images, i.e., is from different spots on the well plate.",
      "votes": null
    },
    {
      "id": "619445",
      "postDate": "09/06/2019 07:58:08",
      "content": "<p>Thanks for sharing useful experiment results!​</p>",
      "rawMarkdown": "Thanks for sharing useful experiment results!​",
      "votes": null
    },
    {
      "id": "625766",
      "postDate": "09/13/2019 12:35:56",
      "content": "<p>I am currently using them as separate images, too. It seems to yield better scores than two inputs to model. I am also considering using L2 of extracted features from both sites as a regularization to the model. Waiting for results</p>",
      "rawMarkdown": "I am currently using them as separate images, too. It seems to yield better scores than two inputs to model. I am also considering using L2 of extracted features from both sites as a regularization to the model. Waiting for results",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 613721,
      "author_name": "interneuron",
      "author_url": "",
      "post_date": "08/30/2019 18:59:29",
      "content": "<p>I've been treating them as straight up separate images, so the training set becomes ~70K images.</p>",
      "votes": null,
      "replies": [
        {
          "id": 615513,
          "author_name": "michelml",
          "author_url": "",
          "post_date": "09/02/2019 03:22:35",
          "content": "<p>Interesting, are you however using both sites to make your predictions on each instance of the test set <a href=\"/interneuron\">@interneuron</a> ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 615542,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "09/02/2019 04:33:22",
          "content": "<p>yes, so total test size is ~39K</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 617690,
          "author_name": "kenkrige",
          "author_url": "",
          "post_date": "09/04/2019 11:48:00",
          "content": "<p>I'm using the 2 sites as separate images. Interestingly, taking the minimum of the 2 predicted probabilities gave me a slightly better LB score than taking the mean. I hadn't even thought of combining them. Maybe something to try, though the image discontinuity at the join may introduce noise that will confuse the model, so I would prefer to leave it as [512, 512, 12] if the two are combined.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 617969,
          "author_name": "micpie",
          "author_url": "",
          "post_date": "09/04/2019 17:08:34",
          "content": "<p>I was also thinking about the same. However, I am not sure if just putting the two 6D site images to 12D images is the best option, as the data is not spatial linked for the two images, i.e., is from different spots on the well plate.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 619445,
      "author_name": "junkoda",
      "author_url": "",
      "post_date": "09/06/2019 07:58:08",
      "content": "<p>Thanks for sharing useful experiment results!​</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 625766,
      "author_name": "roguekk007",
      "author_url": "",
      "post_date": "09/13/2019 12:35:56",
      "content": "<p>I am currently using them as separate images, too. It seems to yield better scores than two inputs to model. I am also considering using L2 of extracted features from both sites as a regularization to the model. Waiting for results</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "613675": "I tried three different approaches to using the images from site1 and site2:\n* combine them into 512x512 image with 12 channels\n* combine into 1024x512 image with 6 channels\n* use them separately as 512x512 images with 6 channels and apply model twice. Like\n```\ndef forward(x):\n  return self.model(x[0]) + self.model(x[1])\n```\n\nWithout changing the model (Resnet18), using 12 channels resulted in worse accuracy. Presumably because only the first conv sees the extra data.\n\nThe later two perform about the same, but separating them seems preferable to me, since I can use other model architectures and share weights more easily between the two sites.\n\nAnyone else found a good way to combine the images from the two sites?",
    "613721": "I've been treating them as straight up separate images, so the training set becomes ~70K images.",
    "615513": "Interesting, are you however using both sites to make your predictions on each instance of the test set @interneuron ?",
    "615542": "yes, so total test size is ~39K",
    "617690": "I'm using the 2 sites as separate images. Interestingly, taking the minimum of the 2 predicted probabilities gave me a slightly better LB score than taking the mean. I hadn't even thought of combining them. Maybe something to try, though the image discontinuity at the join may introduce noise that will confuse the model, so I would prefer to leave it as [512, 512, 12] if the two are combined.",
    "617969": "I was also thinking about the same. However, I am not sure if just putting the two 6D site images to 12D images is the best option, as the data is not spatial linked for the two images, i.e., is from different spots on the well plate.",
    "619445": "Thanks for sharing useful experiment results!​",
    "625766": "I am currently using them as separate images, too. It seems to yield better scores than two inputs to model. I am also considering using L2 of extracted features from both sites as a regularization to the model. Waiting for results"
  },
  "source": "meta"
}