{
  "id": 348783,
  "title": "Question about the HuBMAP images size range",
  "url": "/competitions/hubmap-organ-segmentation/discussion/348783",
  "author_name": "",
  "post_date": "2022-08-29T23:10:58.710653400Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<blockquote>\n  <p>The HuBMAP images range in size from 4500x4500 down to 160x160 pixels.</p>\n</blockquote>\n<p>Is there any distribution ratio for the range in the test set, how common is the upper band version the lower, what's the most used, if there is some statistics to take part of?</p>",
  "messages": [
    {
      "id": "1918858",
      "postDate": "08/29/2022 23:10:58",
      "content": "<blockquote>\n  <p>The HuBMAP images range in size from 4500x4500 down to 160x160 pixels.</p>\n</blockquote>\n<p>Is there any distribution ratio for the range in the test set, how common is the upper band version the lower, what's the most used, if there is some statistics to take part of?</p>",
      "rawMarkdown": "> The HuBMAP images range in size from 4500x4500 down to 160x160 pixels.\n\nIs there any distribution ratio for the range in the test set, how common is the upper band version the lower, what's the most used, if there is some statistics to take part of?",
      "votes": null
    },
    {
      "id": "1919214",
      "postDate": "08/30/2022 08:07:10",
      "content": "<p>What is the source for this statement:</p>\n<pre><code>`The HuBMAP images range in size from 4500x4500 down to 160x160 pixels.`\n</code></pre>\n<p>Training samples range from 2308x2308 to 3070x3070, where the vast majority, 326/351, are sized 3000x3000.</p>\n<p>It would surprise me if the test set would contains 160x160 images, as the smallest train sample is sized 2308x2308.</p>",
      "rawMarkdown": "What is the source for this statement:\n\n    `The HuBMAP images range in size from 4500x4500 down to 160x160 pixels.`\n\nTraining samples range from 2308x2308 to 3070x3070, where the vast majority, 326/351, are sized 3000x3000.\n\nIt would surprise me if the test set would contains 160x160 images, as the smallest train sample is sized 2308x2308.",
      "votes": null
    },
    {
      "id": "1919238",
      "postDate": "08/30/2022 08:32:43",
      "content": "<p>In the data section, in the info about [train/test]_images/<br>\nI think that is the tricky here, big differences in images size between train and test, hence the question :)</p>",
      "rawMarkdown": "In the data section, in the info about [train/test]_images/\nI think that is the tricky here, big differences in images size between train and test, hence the question :)",
      "votes": null
    },
    {
      "id": "1919256",
      "postDate": "08/30/2022 08:59:54",
      "content": "<p>Thanks, I think the differences between train and test data is the main challenge in this competition. The secrecy around the test set, for example with the resolution as you pointed out, makes performance evaluation quite difficult.</p>\n<p><code>The training dataset consists of data from public HPA data, the public test set is a combination of private HPA data and HuBMAP data, and the private test set contains only HuBMAP data</code></p>\n<p>This will likely cause a huge shakeup for the private LB.</p>",
      "rawMarkdown": "Thanks, I think the differences between train and test data is the main challenge in this competition. The secrecy around the test set, for example with the resolution as you pointed out, makes performance evaluation quite difficult.\n\n`The training dataset consists of data from public HPA data, the public test set is a combination of private HPA data and HuBMAP data, and the private test set contains only HuBMAP data`\n\nThis will likely cause a huge shakeup for the private LB.",
      "votes": null
    },
    {
      "id": "1919272",
      "postDate": "08/30/2022 09:24:00",
      "content": "<p>Yes, indeed. The asking for the statistics is if there is a common ratio within in the size range, in the field, if it's the case we can maybe help creating a better model with that knowledge, but if it's random/unknown, that info will don't help, then a generalized model is quite better.</p>",
      "rawMarkdown": "Yes, indeed. The asking for the statistics is if there is a common ratio within in the size range, in the field, if it's the case we can maybe help creating a better model with that knowledge, but if it's random/unknown, that info will don't help, then a generalized model is quite better.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1919214,
      "author_name": "markwijkhuizen",
      "author_url": "",
      "post_date": "08/30/2022 08:07:10",
      "content": "<p>What is the source for this statement:</p>\n<pre><code>`The HuBMAP images range in size from 4500x4500 down to 160x160 pixels.`\n</code></pre>\n<p>Training samples range from 2308x2308 to 3070x3070, where the vast majority, 326/351, are sized 3000x3000.</p>\n<p>It would surprise me if the test set would contains 160x160 images, as the smallest train sample is sized 2308x2308.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1919238,
          "author_name": "kirderf",
          "author_url": "",
          "post_date": "08/30/2022 08:32:43",
          "content": "<p>In the data section, in the info about [train/test]_images/<br>\nI think that is the tricky here, big differences in images size between train and test, hence the question :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1919256,
          "author_name": "markwijkhuizen",
          "author_url": "",
          "post_date": "08/30/2022 08:59:54",
          "content": "<p>Thanks, I think the differences between train and test data is the main challenge in this competition. The secrecy around the test set, for example with the resolution as you pointed out, makes performance evaluation quite difficult.</p>\n<p><code>The training dataset consists of data from public HPA data, the public test set is a combination of private HPA data and HuBMAP data, and the private test set contains only HuBMAP data</code></p>\n<p>This will likely cause a huge shakeup for the private LB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1919272,
          "author_name": "kirderf",
          "author_url": "",
          "post_date": "08/30/2022 09:24:00",
          "content": "<p>Yes, indeed. The asking for the statistics is if there is a common ratio within in the size range, in the field, if it's the case we can maybe help creating a better model with that knowledge, but if it's random/unknown, that info will don't help, then a generalized model is quite better.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1918858": "> The HuBMAP images range in size from 4500x4500 down to 160x160 pixels.\n\nIs there any distribution ratio for the range in the test set, how common is the upper band version the lower, what's the most used, if there is some statistics to take part of?",
    "1919214": "What is the source for this statement:\n\n    `The HuBMAP images range in size from 4500x4500 down to 160x160 pixels.`\n\nTraining samples range from 2308x2308 to 3070x3070, where the vast majority, 326/351, are sized 3000x3000.\n\nIt would surprise me if the test set would contains 160x160 images, as the smallest train sample is sized 2308x2308.",
    "1919238": "In the data section, in the info about [train/test]_images/\nI think that is the tricky here, big differences in images size between train and test, hence the question :)",
    "1919256": "Thanks, I think the differences between train and test data is the main challenge in this competition. The secrecy around the test set, for example with the resolution as you pointed out, makes performance evaluation quite difficult.\n\n`The training dataset consists of data from public HPA data, the public test set is a combination of private HPA data and HuBMAP data, and the private test set contains only HuBMAP data`\n\nThis will likely cause a huge shakeup for the private LB.",
    "1919272": "Yes, indeed. The asking for the statistics is if there is a common ratio within in the size range, in the field, if it's the case we can maybe help creating a better model with that knowledge, but if it's random/unknown, that info will don't help, then a generalized model is quite better."
  },
  "source": "meta"
}