{
  "id": 156713,
  "title": "Size of each image - 2 groups",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/156713",
  "author_name": "",
  "post_date": "2020-06-07T12:30:43.430668700Z",
  "votes": 12,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I calculated the width and height of each image and thought it might be useful to others.</p>\n\n<p>For both width and height there seems to be some quite distinct groups. I am wondering if having different transform functions for these groups might be more beneficial. Resizing a 1000x1000 image to 768x768 isn't going to lose too much detail. However, resizing a 4000x6000 image to 768x768 may do.</p>\n\n<p>Maybe something more along the lines of:\nResize to 768x768 for images originally 3000x3000 or smaller\nResize to 3000x3000 for images originally greater than 3000x3000.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4714993%2F59b79c0f38e10c96fdc1bf52b9ab67b2%2Fheight.pnd.png?generation=1591533072911366&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4714993%2F5dfa3a173a2865e21ef22c68cdb21677%2Fwidth.png?generation=1591532976803249&amp;alt=media\" alt=\"\"></p>\n\n<p>Notebook for those interested: <a href=\"https://www.kaggle.com/blueturtle/siim-list-comprehension-image-sizes/\">https://www.kaggle.com/blueturtle/siim-list-comprehension-image-sizes/</a></p>\n\n<p>I have attached the augmentated csv too with the height and width to save the roughly 4 hours it took the notebook to run :/ </p>\n\n<p>This would require running a number of models equal to the number of splits though on a subset of the data and piecing the results together to form one final submission csv.</p>",
  "messages": [
    {
      "id": "877235",
      "postDate": "06/07/2020 12:30:43",
      "content": "<p>I calculated the width and height of each image and thought it might be useful to others.</p>\n\n<p>For both width and height there seems to be some quite distinct groups. I am wondering if having different transform functions for these groups might be more beneficial. Resizing a 1000x1000 image to 768x768 isn't going to lose too much detail. However, resizing a 4000x6000 image to 768x768 may do.</p>\n\n<p>Maybe something more along the lines of:\nResize to 768x768 for images originally 3000x3000 or smaller\nResize to 3000x3000 for images originally greater than 3000x3000.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4714993%2F59b79c0f38e10c96fdc1bf52b9ab67b2%2Fheight.pnd.png?generation=1591533072911366&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4714993%2F5dfa3a173a2865e21ef22c68cdb21677%2Fwidth.png?generation=1591532976803249&amp;alt=media\" alt=\"\"></p>\n\n<p>Notebook for those interested: <a href=\"https://www.kaggle.com/blueturtle/siim-list-comprehension-image-sizes/\">https://www.kaggle.com/blueturtle/siim-list-comprehension-image-sizes/</a></p>\n\n<p>I have attached the augmentated csv too with the height and width to save the roughly 4 hours it took the notebook to run :/ </p>\n\n<p>This would require running a number of models equal to the number of splits though on a subset of the data and piecing the results together to form one final submission csv.</p>",
      "rawMarkdown": "I calculated the width and height of each image and thought it might be useful to others.\n\nFor both width and height there seems to be some quite distinct groups. I am wondering if having different transform functions for these groups might be more beneficial. Resizing a 1000x1000 image to 768x768 isn't going to lose too much detail. However, resizing a 4000x6000 image to 768x768 may do.\n\nMaybe something more along the lines of:\nResize to 768x768 for images originally 3000x3000 or smaller\nResize to 3000x3000 for images originally greater than 3000x3000.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4714993%2F59b79c0f38e10c96fdc1bf52b9ab67b2%2Fheight.pnd.png?generation=1591533072911366&amp;alt=media)\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4714993%2F5dfa3a173a2865e21ef22c68cdb21677%2Fwidth.png?generation=1591532976803249&amp;alt=media)\n\nNotebook for those interested: https://www.kaggle.com/blueturtle/siim-list-comprehension-image-sizes/\n\nI have attached the augmentated csv too with the height and width to save the roughly 4 hours it took the notebook to run :/ \n\nThis would require running a number of models equal to the number of splits though on a subset of the data and piecing the results together to form one final submission csv.",
      "votes": null
    },
    {
      "id": "877299",
      "postDate": "06/07/2020 13:29:55",
      "content": "<p>Great analysis. Thank you for this work. Processing different image sizes differently may be a good idea.</p>",
      "rawMarkdown": "Great analysis. Thank you for this work. Processing different image sizes differently may be a good idea.",
      "votes": null
    },
    {
      "id": "877633",
      "postDate": "06/07/2020 19:30:20",
      "content": "<p>great idea, would love to give it a try.</p>",
      "rawMarkdown": "great idea, would love to give it a try.",
      "votes": null
    },
    {
      "id": "882464",
      "postDate": "06/11/2020 20:19:00",
      "content": "<p>Nice. I guess this is how the leak is searched ;)</p>",
      "rawMarkdown": "Nice. I guess this is how the leak is searched ;)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 877299,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "06/07/2020 13:29:55",
      "content": "<p>Great analysis. Thank you for this work. Processing different image sizes differently may be a good idea.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 877633,
      "author_name": "rohitsingh9990",
      "author_url": "",
      "post_date": "06/07/2020 19:30:20",
      "content": "<p>great idea, would love to give it a try.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 882464,
      "author_name": "mks2192",
      "author_url": "",
      "post_date": "06/11/2020 20:19:00",
      "content": "<p>Nice. I guess this is how the leak is searched ;)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "877235": "I calculated the width and height of each image and thought it might be useful to others.\n\nFor both width and height there seems to be some quite distinct groups. I am wondering if having different transform functions for these groups might be more beneficial. Resizing a 1000x1000 image to 768x768 isn't going to lose too much detail. However, resizing a 4000x6000 image to 768x768 may do.\n\nMaybe something more along the lines of:\nResize to 768x768 for images originally 3000x3000 or smaller\nResize to 3000x3000 for images originally greater than 3000x3000.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4714993%2F59b79c0f38e10c96fdc1bf52b9ab67b2%2Fheight.pnd.png?generation=1591533072911366&amp;alt=media)\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4714993%2F5dfa3a173a2865e21ef22c68cdb21677%2Fwidth.png?generation=1591532976803249&amp;alt=media)\n\nNotebook for those interested: https://www.kaggle.com/blueturtle/siim-list-comprehension-image-sizes/\n\nI have attached the augmentated csv too with the height and width to save the roughly 4 hours it took the notebook to run :/ \n\nThis would require running a number of models equal to the number of splits though on a subset of the data and piecing the results together to form one final submission csv.",
    "877299": "Great analysis. Thank you for this work. Processing different image sizes differently may be a good idea.",
    "877633": "great idea, would love to give it a try.",
    "882464": "Nice. I guess this is how the leak is searched ;)"
  },
  "source": "meta"
}