{
  "id": 235197,
  "title": "what if drop 75% tiles without mask to use higher resolution",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/235197",
  "author_name": "",
  "post_date": "2021-04-28T08:51:31.919298700Z",
  "votes": 5,
  "comment_count": 6,
  "views": 0,
  "content": "<p>In \"deepflash\"'s work<br>\ndo sample most on masks area and less and less on other \"trivial\" area. seems not hurt performance.</p>\n<p>I reduce pictures 2 or 4 time to get smaller file size <br>\ndue to training time<br>\neven in TPU <br>\nonly 10 tiffs 512x512 take about 5 hours to training.<br>\nwhile 40 tiffs 256x256 take about 2 hours</p>\n<p>what if I drop most pics without mask for using 1024 resolution pictures<br>\nless pictures but higher resolution <br>\nwith keep in all masks.</p>\n<p>In my intuition <br>\nhigher resolution bring much more information and detail will improve precise<br>\n“trivial area” seems give no benefit</p>\n<p>I read some paper about efficientnet and Unet , but seems not so many talk about \"trivial area pics\" give benefit or hurt </p>\n<p>what do you think ? <br>\nsomeone have some experiments?</p>",
  "messages": [
    {
      "id": "1286649",
      "postDate": "04/28/2021 08:51:31",
      "content": "<p>In \"deepflash\"'s work<br>\ndo sample most on masks area and less and less on other \"trivial\" area. seems not hurt performance.</p>\n<p>I reduce pictures 2 or 4 time to get smaller file size <br>\ndue to training time<br>\neven in TPU <br>\nonly 10 tiffs 512x512 take about 5 hours to training.<br>\nwhile 40 tiffs 256x256 take about 2 hours</p>\n<p>what if I drop most pics without mask for using 1024 resolution pictures<br>\nless pictures but higher resolution <br>\nwith keep in all masks.</p>\n<p>In my intuition <br>\nhigher resolution bring much more information and detail will improve precise<br>\n“trivial area” seems give no benefit</p>\n<p>I read some paper about efficientnet and Unet , but seems not so many talk about \"trivial area pics\" give benefit or hurt </p>\n<p>what do you think ? <br>\nsomeone have some experiments?</p>",
      "rawMarkdown": "In \"deepflash\"'s work\ndo sample most on masks area and less and less on other \"trivial\" area. seems not hurt performance.\n\nI reduce pictures 2 or 4 time to get smaller file size \ndue to training time\neven in TPU \nonly 10 tiffs 512x512 take about 5 hours to training.\nwhile 40 tiffs 256x256 take about 2 hours\n\nwhat if I drop most pics without mask for using 1024 resolution pictures\nless pictures but higher resolution \nwith keep in all masks.\n\nIn my intuition \nhigher resolution bring much more information and detail will improve precise\n“trivial area” seems give no benefit\n\nI read some paper about efficientnet and Unet , but seems not so many talk about \"trivial area pics\" give benefit or hurt \n\n\nwhat do you think ? \nsomeone have some experiments?",
      "votes": null
    },
    {
      "id": "1286997",
      "postDate": "04/28/2021 15:47:05",
      "content": "<p>Hi, our idea behind deepflash is, that we want to avoid training on images without any information. Therefore, we upsample (in terms of sampling frequency and not resolution) tiles that contain the target class \"Glomerulus\" while large amounts of the image can be ignored.<br>\nFurthermore, we do sample images with Cortex more often than Medulla or \"Background\", because there are more examples, which can be confused with the target label.</p>\n<p>In our opinion, there is no reason to train the model on many all glass-background or black tiles, since they do look similar and provide no useful information to train the model. This is why we decided to sample those with a small probability.</p>",
      "rawMarkdown": "Hi, our idea behind deepflash is, that we want to avoid training on images without any information. Therefore, we upsample (in terms of sampling frequency and not resolution) tiles that contain the target class \"Glomerulus\" while large amounts of the image can be ignored.\nFurthermore, we do sample images with Cortex more often than Medulla or \"Background\", because there are more examples, which can be confused with the target label.\n\nIn our opinion, there is no reason to train the model on many all glass-background or black tiles, since they do look similar and provide no useful information to train the model. This is why we decided to sample those with a small probability.",
      "votes": null
    },
    {
      "id": "1287289",
      "postDate": "04/28/2021 21:57:19",
      "content": "<p>``You don't have to drop image, since this compition is binary label, you can simply maintain positive/negative mask ratio on the fly. random sample another image if the mask isn't want you want, until you sample a positive label.</p>",
      "rawMarkdown": "``You don't have to drop image, since this compition is binary label, you can simply maintain positive/negative mask ratio on the fly. random sample another image if the mask isn't want you want, until you sample a positive label.",
      "votes": null
    },
    {
      "id": "1287360",
      "postDate": "04/29/2021 01:54:56",
      "content": "<p>does it need a classifier head to know which is positive, which is not? I don't know if segmentation head alone can handle  inference well.</p>",
      "rawMarkdown": "does it need a classifier head to know which is positive, which is not? I don't know if segmentation head alone can handle  inference well.",
      "votes": null
    },
    {
      "id": "1287515",
      "postDate": "04/29/2021 06:31:34",
      "content": "<p>Does it mean a fine drop tiles sampling strategies, rather than coarse drop tiles by random?<br>\nIn your opnion</p>",
      "rawMarkdown": "Does it mean a fine drop tiles sampling strategies, rather than coarse drop tiles by random?\nIn your opnion",
      "votes": null
    },
    {
      "id": "1287644",
      "postDate": "04/29/2021 09:05:43",
      "content": "<p>I don't think I understand your question.  <br>\nWe don't drop the tiles, we just sample them by random with different sampling probabilities:<br>\nGlomeruli &gt; Cortex &gt; Medulla &gt;&gt; Rest</p>",
      "rawMarkdown": "I don't think I understand your question.  \nWe don't drop the tiles, we just sample them by random with different sampling probabilities:\nGlomeruli > Cortex > Medulla >> Rest",
      "votes": null
    },
    {
      "id": "1287757",
      "postDate": "04/29/2021 11:01:17",
      "content": "<p>I would say it won't help much:</p>\n<ol>\n<li>it only two class</li>\n<li>the boundary effect (positive cells be cut off on the edge), create lots of noise on the boundary. this problem will propagate back to your model, which is not an desired error you want to learn.</li>\n</ol>",
      "rawMarkdown": "I would say it won't help much:\n1.  it only two class\n2. the boundary effect (positive cells be cut off on the edge), create lots of noise on the boundary. this problem will propagate back to your model, which is not an desired error you want to learn.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1286997,
      "author_name": "theudas",
      "author_url": "",
      "post_date": "04/28/2021 15:47:05",
      "content": "<p>Hi, our idea behind deepflash is, that we want to avoid training on images without any information. Therefore, we upsample (in terms of sampling frequency and not resolution) tiles that contain the target class \"Glomerulus\" while large amounts of the image can be ignored.<br>\nFurthermore, we do sample images with Cortex more often than Medulla or \"Background\", because there are more examples, which can be confused with the target label.</p>\n<p>In our opinion, there is no reason to train the model on many all glass-background or black tiles, since they do look similar and provide no useful information to train the model. This is why we decided to sample those with a small probability.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1287515,
          "author_name": "drzhuzhe",
          "author_url": "",
          "post_date": "04/29/2021 06:31:34",
          "content": "<p>Does it mean a fine drop tiles sampling strategies, rather than coarse drop tiles by random?<br>\nIn your opnion</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1287644,
          "author_name": "theudas",
          "author_url": "",
          "post_date": "04/29/2021 09:05:43",
          "content": "<p>I don't think I understand your question.  <br>\nWe don't drop the tiles, we just sample them by random with different sampling probabilities:<br>\nGlomeruli &gt; Cortex &gt; Medulla &gt;&gt; Rest</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1287289,
      "author_name": "gody7334",
      "author_url": "",
      "post_date": "04/28/2021 21:57:19",
      "content": "<p>``You don't have to drop image, since this compition is binary label, you can simply maintain positive/negative mask ratio on the fly. random sample another image if the mask isn't want you want, until you sample a positive label.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1287360,
      "author_name": "plugin1689",
      "author_url": "",
      "post_date": "04/29/2021 01:54:56",
      "content": "<p>does it need a classifier head to know which is positive, which is not? I don't know if segmentation head alone can handle  inference well.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1287757,
          "author_name": "gody7334",
          "author_url": "",
          "post_date": "04/29/2021 11:01:17",
          "content": "<p>I would say it won't help much:</p>\n<ol>\n<li>it only two class</li>\n<li>the boundary effect (positive cells be cut off on the edge), create lots of noise on the boundary. this problem will propagate back to your model, which is not an desired error you want to learn.</li>\n</ol>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1286649": "In \"deepflash\"'s work\ndo sample most on masks area and less and less on other \"trivial\" area. seems not hurt performance.\n\nI reduce pictures 2 or 4 time to get smaller file size \ndue to training time\neven in TPU \nonly 10 tiffs 512x512 take about 5 hours to training.\nwhile 40 tiffs 256x256 take about 2 hours\n\nwhat if I drop most pics without mask for using 1024 resolution pictures\nless pictures but higher resolution \nwith keep in all masks.\n\nIn my intuition \nhigher resolution bring much more information and detail will improve precise\n“trivial area” seems give no benefit\n\nI read some paper about efficientnet and Unet , but seems not so many talk about \"trivial area pics\" give benefit or hurt \n\n\nwhat do you think ? \nsomeone have some experiments?",
    "1286997": "Hi, our idea behind deepflash is, that we want to avoid training on images without any information. Therefore, we upsample (in terms of sampling frequency and not resolution) tiles that contain the target class \"Glomerulus\" while large amounts of the image can be ignored.\nFurthermore, we do sample images with Cortex more often than Medulla or \"Background\", because there are more examples, which can be confused with the target label.\n\nIn our opinion, there is no reason to train the model on many all glass-background or black tiles, since they do look similar and provide no useful information to train the model. This is why we decided to sample those with a small probability.",
    "1287289": "``You don't have to drop image, since this compition is binary label, you can simply maintain positive/negative mask ratio on the fly. random sample another image if the mask isn't want you want, until you sample a positive label.",
    "1287360": "does it need a classifier head to know which is positive, which is not? I don't know if segmentation head alone can handle  inference well.",
    "1287515": "Does it mean a fine drop tiles sampling strategies, rather than coarse drop tiles by random?\nIn your opnion",
    "1287644": "I don't think I understand your question.  \nWe don't drop the tiles, we just sample them by random with different sampling probabilities:\nGlomeruli > Cortex > Medulla >> Rest",
    "1287757": "I would say it won't help much:\n1.  it only two class\n2. the boundary effect (positive cells be cut off on the edge), create lots of noise on the boundary. this problem will propagate back to your model, which is not an desired error you want to learn."
  },
  "source": "meta"
}