{
  "id": 199158,
  "title": "data size and manipulation",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/199158",
  "author_name": "",
  "post_date": "2020-11-24T16:48:24.184880800Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>hi,<br>\nI wanted to know that image sizes are very large to be trained but if the size of the images are to be reduced to 256x256, wouldn't be there loss of information and result in less accurate results.</p>\n<p>What if each and images are split up into n times(but finite obviously) so that every glomerulli pixels can also become available in my dataset.</p>\n<p>Please kindly guide me, any help will be appreciated</p>\n<p>Thank you</p>",
  "messages": [
    {
      "id": "1089646",
      "postDate": "11/24/2020 16:48:24",
      "content": "<p>hi,<br>\nI wanted to know that image sizes are very large to be trained but if the size of the images are to be reduced to 256x256, wouldn't be there loss of information and result in less accurate results.</p>\n<p>What if each and images are split up into n times(but finite obviously) so that every glomerulli pixels can also become available in my dataset.</p>\n<p>Please kindly guide me, any help will be appreciated</p>\n<p>Thank you</p>",
      "rawMarkdown": "hi,\nI wanted to know that image sizes are very large to be trained but if the size of the images are to be reduced to 256x256, wouldn't be there loss of information and result in less accurate results.\n\nWhat if each and images are split up into n times(but finite obviously) so that every glomerulli pixels can also become available in my dataset.\n\nPlease kindly guide me, any help will be appreciated\n\nThank you",
      "votes": null
    },
    {
      "id": "1089717",
      "postDate": "11/24/2020 18:12:10",
      "content": "<p>There are a couple of strategies when training off of large images but in the end, it comes down to the fact that not every pixel of the image <em>matters</em>.  Think, do you need 4k footage to identify a person or Here are some examples of what can be done:</p>\n<ul>\n<li>You can preprocess the images to reduce the amount of information to process (this could be resizing but other strategies exist)</li>\n<li>Run the images through a pre-trained network to then train on features versus raw pixel values.</li>\n</ul>\n<p>I think you're right in wondering if there will be too much loss of information but depending on the context this loss of information might be alright. I think splitting the images into multiple parts is an interesting strategy. Assuming that we can determine the correct output by parts of the image instead of having to rely on the full image. This seems to be true based on this exploratory notebook: <a href=\"https://www.kaggle.com/ihelon/hubmap-exploratory-data-analysis\" target=\"_blank\">https://www.kaggle.com/ihelon/hubmap-exploratory-data-analysis</a></p>\n<p>So I think this seems like a decent idea to split the single image into many examples. The only thing is you'll likely have to figure out is creating appropriate outputs for your sliced-up images. But I think this seems doable.</p>\n<p>Hope that's a bit helpful and good luck!</p>\n<blockquote>\n  <p>PS: You might want to check out this notebook that plays around with using 256x256 images to prototype the model: <a href=\"https://www.kaggle.com/iafoss/256x256-images\" target=\"_blank\">https://www.kaggle.com/iafoss/256x256-images</a></p>\n</blockquote>",
      "rawMarkdown": "There are a couple of strategies when training off of large images but in the end, it comes down to the fact that not every pixel of the image _matters_.  Think, do you need 4k footage to identify a person or Here are some examples of what can be done:\n\n- You can preprocess the images to reduce the amount of information to process (this could be resizing but other strategies exist)\n- Run the images through a pre-trained network to then train on features versus raw pixel values.\n\nI think you're right in wondering if there will be too much loss of information but depending on the context this loss of information might be alright. I think splitting the images into multiple parts is an interesting strategy. Assuming that we can determine the correct output by parts of the image instead of having to rely on the full image. This seems to be true based on this exploratory notebook: https://www.kaggle.com/ihelon/hubmap-exploratory-data-analysis\n\nSo I think this seems like a decent idea to split the single image into many examples. The only thing is you'll likely have to figure out is creating appropriate outputs for your sliced-up images. But I think this seems doable.\n\nHope that's a bit helpful and good luck!\n\n>PS: You might want to check out this notebook that plays around with using 256x256 images to prototype the model: https://www.kaggle.com/iafoss/256x256-images",
      "votes": null
    },
    {
      "id": "1090292",
      "postDate": "11/25/2020 08:32:12",
      "content": "<p>thanks friend</p>",
      "rawMarkdown": "thanks friend",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1089717,
      "author_name": "mrgeislinger",
      "author_url": "",
      "post_date": "11/24/2020 18:12:10",
      "content": "<p>There are a couple of strategies when training off of large images but in the end, it comes down to the fact that not every pixel of the image <em>matters</em>.  Think, do you need 4k footage to identify a person or Here are some examples of what can be done:</p>\n<ul>\n<li>You can preprocess the images to reduce the amount of information to process (this could be resizing but other strategies exist)</li>\n<li>Run the images through a pre-trained network to then train on features versus raw pixel values.</li>\n</ul>\n<p>I think you're right in wondering if there will be too much loss of information but depending on the context this loss of information might be alright. I think splitting the images into multiple parts is an interesting strategy. Assuming that we can determine the correct output by parts of the image instead of having to rely on the full image. This seems to be true based on this exploratory notebook: <a href=\"https://www.kaggle.com/ihelon/hubmap-exploratory-data-analysis\" target=\"_blank\">https://www.kaggle.com/ihelon/hubmap-exploratory-data-analysis</a></p>\n<p>So I think this seems like a decent idea to split the single image into many examples. The only thing is you'll likely have to figure out is creating appropriate outputs for your sliced-up images. But I think this seems doable.</p>\n<p>Hope that's a bit helpful and good luck!</p>\n<blockquote>\n  <p>PS: You might want to check out this notebook that plays around with using 256x256 images to prototype the model: <a href=\"https://www.kaggle.com/iafoss/256x256-images\" target=\"_blank\">https://www.kaggle.com/iafoss/256x256-images</a></p>\n</blockquote>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1090292,
      "author_name": "roshaanzafar",
      "author_url": "",
      "post_date": "11/25/2020 08:32:12",
      "content": "<p>thanks friend</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1089646": "hi,\nI wanted to know that image sizes are very large to be trained but if the size of the images are to be reduced to 256x256, wouldn't be there loss of information and result in less accurate results.\n\nWhat if each and images are split up into n times(but finite obviously) so that every glomerulli pixels can also become available in my dataset.\n\nPlease kindly guide me, any help will be appreciated\n\nThank you",
    "1089717": "There are a couple of strategies when training off of large images but in the end, it comes down to the fact that not every pixel of the image _matters_.  Think, do you need 4k footage to identify a person or Here are some examples of what can be done:\n\n- You can preprocess the images to reduce the amount of information to process (this could be resizing but other strategies exist)\n- Run the images through a pre-trained network to then train on features versus raw pixel values.\n\nI think you're right in wondering if there will be too much loss of information but depending on the context this loss of information might be alright. I think splitting the images into multiple parts is an interesting strategy. Assuming that we can determine the correct output by parts of the image instead of having to rely on the full image. This seems to be true based on this exploratory notebook: https://www.kaggle.com/ihelon/hubmap-exploratory-data-analysis\n\nSo I think this seems like a decent idea to split the single image into many examples. The only thing is you'll likely have to figure out is creating appropriate outputs for your sliced-up images. But I think this seems doable.\n\nHope that's a bit helpful and good luck!\n\n>PS: You might want to check out this notebook that plays around with using 256x256 images to prototype the model: https://www.kaggle.com/iafoss/256x256-images",
    "1090292": "thanks friend"
  },
  "source": "meta"
}