{
  "id": 94521,
  "title": "Demand for re-sized data-sets?",
  "url": "/competitions/open-images-2019-object-detection/discussion/94521",
  "author_name": "",
  "post_date": "2019-06-05T02:08:34.633040400Z",
  "votes": 29,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Hey folks!</p>\n\n<p>Noticed this is one of the many competitions that has hundreds of gigabytes of data to process for a submission. I recently took part in another one of this sort (Google Landmark Recognition), and was wondering if there is demand for a re-sized data-set (think 256x256 format compressed images).</p>\n\n<p>If so, let me know, and I'll get started on re-sizing them.</p>",
  "messages": [
    {
      "id": "543973",
      "postDate": "06/05/2019 02:08:34",
      "content": "<p>Hey folks!</p>\n\n<p>Noticed this is one of the many competitions that has hundreds of gigabytes of data to process for a submission. I recently took part in another one of this sort (Google Landmark Recognition), and was wondering if there is demand for a re-sized data-set (think 256x256 format compressed images).</p>\n\n<p>If so, let me know, and I'll get started on re-sizing them.</p>",
      "rawMarkdown": "Hey folks!\n\nNoticed this is one of the many competitions that has hundreds of gigabytes of data to process for a submission. I recently took part in another one of this sort (Google Landmark Recognition), and was wondering if there is demand for a re-sized data-set (think 256x256 format compressed images).\n\nIf so, let me know, and I'll get started on re-sizing them.",
      "votes": null
    },
    {
      "id": "543981",
      "postDate": "06/05/2019 02:14:58",
      "content": "<p>Yes! I would very much appreciate it. I’m sure others would also. Thanks !</p>",
      "rawMarkdown": "Yes! I would very much appreciate it. I’m sure others would also. Thanks !",
      "votes": null
    },
    {
      "id": "544004",
      "postDate": "06/05/2019 02:55:59",
      "content": "<p>Yes! really need resized dataset, thank you</p>",
      "rawMarkdown": "Yes! really need resized dataset, thank you",
      "votes": null
    },
    {
      "id": "544051",
      "postDate": "06/05/2019 04:22:55",
      "content": "<p>Please help people out as you did in Google landmark challenge.:)</p>",
      "rawMarkdown": "Please help people out as you did in Google landmark challenge.:)",
      "votes": null
    },
    {
      "id": "544101",
      "postDate": "06/05/2019 05:59:39",
      "content": "<p>+1 in need. Will be very grateful :)</p>",
      "rawMarkdown": "1 in need. Will be very grateful :)",
      "votes": null
    },
    {
      "id": "544242",
      "postDate": "06/05/2019 09:53:32",
      "content": "<p>me too</p>",
      "rawMarkdown": "me too",
      "votes": null
    },
    {
      "id": "544478",
      "postDate": "06/05/2019 14:53:14",
      "content": "<p>There are different ways to downsample an image. Which would you use? Does someone know how much it matters in terms of building models and accuracy of the predcition which algorithm is used?</p>",
      "rawMarkdown": "There are different ways to downsample an image. Which would you use? Does someone know how much it matters in terms of building models and accuracy of the predcition which algorithm is used?",
      "votes": null
    },
    {
      "id": "544494",
      "postDate": "06/05/2019 15:07:20",
      "content": "<p>I use <code>imgp</code> for multi-threaded re-sizing capabilities (progressive quality filter at 90% of base, and do a hard re-size to <code>256x256</code> non-proportional sized JPG files).</p>\n\n<p>This gives you an output of all images at <code>256x256</code> crops with 90% quality. For the last competition, I did proportional re-sizing so that the outputs all had a <code>256px</code> width, but different heights. Multiple people reached out and suggest I do proportional sizing though, since many models take square images as input, and it's easier for the end-users.</p>\n\n<p>Overall, you lose a couple percent of efficiency on the re-size (I recently did the Google Landmark Challenge where I lost about 6-11% additional image data). But, for the average competitor that can't download the hundreds of GB of data, this is the best way to compete.</p>",
      "rawMarkdown": "I use `imgp` for multi-threaded re-sizing capabilities (progressive quality filter at 90% of base, and do a hard re-size to `256x256` non-proportional sized JPG files).\n\nThis gives you an output of all images at `256x256` crops with 90% quality. For the last competition, I did proportional re-sizing so that the outputs all had a `256px` width, but different heights. Multiple people reached out and suggest I do proportional sizing though, since many models take square images as input, and it's easier for the end-users.\n\nOverall, you lose a couple percent of efficiency on the re-size (I recently did the Google Landmark Challenge where I lost about 6-11% additional image data). But, for the average competitor that can't download the hundreds of GB of data, this is the best way to compete.",
      "votes": null
    },
    {
      "id": "544543",
      "postDate": "06/05/2019 16:23:02",
      "content": "<p>Sure! Having a smaller dataset would be great for most users. And that includes me.</p>\n\n<p>I am solely asking, what impact different downsampling methods have. There is aspect ratio as you mentioned. (I would vote for the longer side to be some fixed value, e.g. 256px. One could always fill to a square image with black or gray or so, if one algorithm needs that.) However, also the downsampling itself can be done different ways. One can do simple averages over the original pixels that go into one downsampled one, one can use some weighting with window functions / kernels, one can Fourier transform the image and cut higher frequencies, etc.\nDoes anybody have a grasp on what the impact for modelling is?</p>",
      "rawMarkdown": "Sure! Having a smaller dataset would be great for most users. And that includes me.\n\nI am solely asking, what impact different downsampling methods have. There is aspect ratio as you mentioned. (I would vote for the longer side to be some fixed value, e.g. 256px. One could always fill to a square image with black or gray or so, if one algorithm needs that.) However, also the downsampling itself can be done different ways. One can do simple averages over the original pixels that go into one downsampled one, one can use some weighting with window functions / kernels, one can Fourier transform the image and cut higher frequencies, etc.\nDoes anybody have a grasp on what the impact for modelling is?",
      "votes": null
    },
    {
      "id": "545904",
      "postDate": "06/06/2019 03:00:13",
      "content": "<p>Well I am new to this Image recognition stuff ....Can anyone share a tutorial for this ? </p>",
      "rawMarkdown": "Well I am new to this Image recognition stuff ....Can anyone share a tutorial for this ?",
      "votes": null
    },
    {
      "id": "546004",
      "postDate": "06/06/2019 06:00:14",
      "content": "<p>Hi Anish! Kudos to you for your great work in recognition. Although I didn't use your datasets last time since I was using Kernels solely and used the tar files in a download-train-delete kind of way. But here in this competition, the tar files are each of ~30 GB, which makes them impossible to use in kernels. If you happen to go forward and resize the train set, can you break them into tar files such that the uncompressed sizes of each of those are around 2-3 GB at max? In that way, a lot more people like me will be able to participate who otherwise have no such resources.</p>",
      "rawMarkdown": "Hi Anish! Kudos to you for your great work in recognition. Although I didn't use your datasets last time since I was using Kernels solely and used the tar files in a download-train-delete kind of way. But here in this competition, the tar files are each of ~30 GB, which makes them impossible to use in kernels. If you happen to go forward and resize the train set, can you break them into tar files such that the uncompressed sizes of each of those are around 2-3 GB at max? In that way, a lot more people like me will be able to participate who otherwise have no such resources.",
      "votes": null
    },
    {
      "id": "546239",
      "postDate": "06/06/2019 11:35:34",
      "content": "<p>Hey <a href=\"/mayukh18\">@mayukh18</a>!</p>\n\n<p>For sure. I've already shared the <code>test</code> and <code>validation</code> sets in a compressed format. I'll be re-sizing and then splitting up the <code>train</code> data-set into small zip files, and sharing them through a Kaggle dataset directly.</p>\n\n<p>I'll drop a message in this thread in the next couple of days when the download and re-size is complete on my end.</p>",
      "rawMarkdown": "Hey @mayukh18!\n\nFor sure. I've already shared the `test` and `validation` sets in a compressed format. I'll be re-sizing and then splitting up the `train` data-set into small zip files, and sharing them through a Kaggle dataset directly.\n\nI'll drop a message in this thread in the next couple of days when the download and re-size is complete on my end.",
      "votes": null
    },
    {
      "id": "550600",
      "postDate": "06/11/2019 19:47:14",
      "content": "<p>You are amazing for doing this, thanks!</p>",
      "rawMarkdown": "You are amazing for doing this, thanks!",
      "votes": null
    },
    {
      "id": "552659",
      "postDate": "06/14/2019 09:42:09",
      "content": "<blockquote>\n  <p>Multiple people reached out and suggest I do proportional sizing though, since many models take square images as input, and it's easier for the end-users.</p>\n</blockquote>\n\n<p>I'd be interested to know more about the possibilities and the tradeoffs when resizing images and changing aspect ratios/filling up images. Could you kindly point to those comments you received and/or to some comprehensive resource on the topic? Many thanks in advance!</p>",
      "rawMarkdown": "&gt; Multiple people reached out and suggest I do proportional sizing though, since many models take square images as input, and it's easier for the end-users.\n\nI'd be interested to know more about the possibilities and the tradeoffs when resizing images and changing aspect ratios/filling up images. Could you kindly point to those comments you received and/or to some comprehensive resource on the topic? Many thanks in advance!",
      "votes": null
    },
    {
      "id": "555052",
      "postDate": "06/18/2019 11:23:44",
      "content": "<p><a href=\"/anishagnihotri\">@anishagnihotri</a> I am still not clear about the \"resizing\" thing. Could you point me to article/paper where the trade-off between size and quality is clearly explained.</p>\n\n<p>right now i am skimming through <a href=\"https://github.com/jarun/imgp\">https://github.com/jarun/imgp</a></p>",
      "rawMarkdown": "anishagnihotri I am still not clear about the \"resizing\" thing. Could you point me to article/paper where the trade-off between size and quality is clearly explained.\n\nright now i am skimming through https://github.com/jarun/imgp",
      "votes": null
    },
    {
      "id": "555054",
      "postDate": "06/18/2019 11:25:44",
      "content": "<p>just try to read through kaggle kernels and try googling/ stackoverflow for queries. There are tons of tutorials out there on computer vision.</p>",
      "rawMarkdown": "just try to read through kaggle kernels and try googling/ stackoverflow for queries. There are tons of tutorials out there on computer vision.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 543981,
      "author_name": "interneuron",
      "author_url": "",
      "post_date": "06/05/2019 02:14:58",
      "content": "<p>Yes! I would very much appreciate it. I’m sure others would also. Thanks !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 544004,
      "author_name": "cswwp347724",
      "author_url": "",
      "post_date": "06/05/2019 02:55:59",
      "content": "<p>Yes! really need resized dataset, thank you</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 544051,
      "author_name": "chandanverma",
      "author_url": "",
      "post_date": "06/05/2019 04:22:55",
      "content": "<p>Please help people out as you did in Google landmark challenge.:)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 544101,
      "author_name": "volhaleusha",
      "author_url": "",
      "post_date": "06/05/2019 05:59:39",
      "content": "<p>+1 in need. Will be very grateful :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 544242,
      "author_name": "validmodel",
      "author_url": "",
      "post_date": "06/05/2019 09:53:32",
      "content": "<p>me too</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 544478,
      "author_name": "bernir",
      "author_url": "",
      "post_date": "06/05/2019 14:53:14",
      "content": "<p>There are different ways to downsample an image. Which would you use? Does someone know how much it matters in terms of building models and accuracy of the predcition which algorithm is used?</p>",
      "votes": null,
      "replies": [
        {
          "id": 544494,
          "author_name": "anishagnihotri",
          "author_url": "",
          "post_date": "06/05/2019 15:07:20",
          "content": "<p>I use <code>imgp</code> for multi-threaded re-sizing capabilities (progressive quality filter at 90% of base, and do a hard re-size to <code>256x256</code> non-proportional sized JPG files).</p>\n\n<p>This gives you an output of all images at <code>256x256</code> crops with 90% quality. For the last competition, I did proportional re-sizing so that the outputs all had a <code>256px</code> width, but different heights. Multiple people reached out and suggest I do proportional sizing though, since many models take square images as input, and it's easier for the end-users.</p>\n\n<p>Overall, you lose a couple percent of efficiency on the re-size (I recently did the Google Landmark Challenge where I lost about 6-11% additional image data). But, for the average competitor that can't download the hundreds of GB of data, this is the best way to compete.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 544543,
          "author_name": "bernir",
          "author_url": "",
          "post_date": "06/05/2019 16:23:02",
          "content": "<p>Sure! Having a smaller dataset would be great for most users. And that includes me.</p>\n\n<p>I am solely asking, what impact different downsampling methods have. There is aspect ratio as you mentioned. (I would vote for the longer side to be some fixed value, e.g. 256px. One could always fill to a square image with black or gray or so, if one algorithm needs that.) However, also the downsampling itself can be done different ways. One can do simple averages over the original pixels that go into one downsampled one, one can use some weighting with window functions / kernels, one can Fourier transform the image and cut higher frequencies, etc.\nDoes anybody have a grasp on what the impact for modelling is?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 552659,
          "author_name": "atineose",
          "author_url": "",
          "post_date": "06/14/2019 09:42:09",
          "content": "<blockquote>\n  <p>Multiple people reached out and suggest I do proportional sizing though, since many models take square images as input, and it's easier for the end-users.</p>\n</blockquote>\n\n<p>I'd be interested to know more about the possibilities and the tradeoffs when resizing images and changing aspect ratios/filling up images. Could you kindly point to those comments you received and/or to some comprehensive resource on the topic? Many thanks in advance!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 555052,
          "author_name": "bakkuuuu",
          "author_url": "",
          "post_date": "06/18/2019 11:23:44",
          "content": "<p><a href=\"/anishagnihotri\">@anishagnihotri</a> I am still not clear about the \"resizing\" thing. Could you point me to article/paper where the trade-off between size and quality is clearly explained.</p>\n\n<p>right now i am skimming through <a href=\"https://github.com/jarun/imgp\">https://github.com/jarun/imgp</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 545904,
      "author_name": "rahhy1",
      "author_url": "",
      "post_date": "06/06/2019 03:00:13",
      "content": "<p>Well I am new to this Image recognition stuff ....Can anyone share a tutorial for this ? </p>",
      "votes": null,
      "replies": [
        {
          "id": 555054,
          "author_name": "bakkuuuu",
          "author_url": "",
          "post_date": "06/18/2019 11:25:44",
          "content": "<p>just try to read through kaggle kernels and try googling/ stackoverflow for queries. There are tons of tutorials out there on computer vision.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 546004,
      "author_name": "mayukh18",
      "author_url": "",
      "post_date": "06/06/2019 06:00:14",
      "content": "<p>Hi Anish! Kudos to you for your great work in recognition. Although I didn't use your datasets last time since I was using Kernels solely and used the tar files in a download-train-delete kind of way. But here in this competition, the tar files are each of ~30 GB, which makes them impossible to use in kernels. If you happen to go forward and resize the train set, can you break them into tar files such that the uncompressed sizes of each of those are around 2-3 GB at max? In that way, a lot more people like me will be able to participate who otherwise have no such resources.</p>",
      "votes": null,
      "replies": [
        {
          "id": 546239,
          "author_name": "anishagnihotri",
          "author_url": "",
          "post_date": "06/06/2019 11:35:34",
          "content": "<p>Hey <a href=\"/mayukh18\">@mayukh18</a>!</p>\n\n<p>For sure. I've already shared the <code>test</code> and <code>validation</code> sets in a compressed format. I'll be re-sizing and then splitting up the <code>train</code> data-set into small zip files, and sharing them through a Kaggle dataset directly.</p>\n\n<p>I'll drop a message in this thread in the next couple of days when the download and re-size is complete on my end.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 550600,
      "author_name": "jamesrequa",
      "author_url": "",
      "post_date": "06/11/2019 19:47:14",
      "content": "<p>You are amazing for doing this, thanks!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "543973": "Hey folks!\n\nNoticed this is one of the many competitions that has hundreds of gigabytes of data to process for a submission. I recently took part in another one of this sort (Google Landmark Recognition), and was wondering if there is demand for a re-sized data-set (think 256x256 format compressed images).\n\nIf so, let me know, and I'll get started on re-sizing them.",
    "543981": "Yes! I would very much appreciate it. I’m sure others would also. Thanks !",
    "544004": "Yes! really need resized dataset, thank you",
    "544051": "Please help people out as you did in Google landmark challenge.:)",
    "544101": "1 in need. Will be very grateful :)",
    "544242": "me too",
    "544478": "There are different ways to downsample an image. Which would you use? Does someone know how much it matters in terms of building models and accuracy of the predcition which algorithm is used?",
    "544494": "I use `imgp` for multi-threaded re-sizing capabilities (progressive quality filter at 90% of base, and do a hard re-size to `256x256` non-proportional sized JPG files).\n\nThis gives you an output of all images at `256x256` crops with 90% quality. For the last competition, I did proportional re-sizing so that the outputs all had a `256px` width, but different heights. Multiple people reached out and suggest I do proportional sizing though, since many models take square images as input, and it's easier for the end-users.\n\nOverall, you lose a couple percent of efficiency on the re-size (I recently did the Google Landmark Challenge where I lost about 6-11% additional image data). But, for the average competitor that can't download the hundreds of GB of data, this is the best way to compete.",
    "544543": "Sure! Having a smaller dataset would be great for most users. And that includes me.\n\nI am solely asking, what impact different downsampling methods have. There is aspect ratio as you mentioned. (I would vote for the longer side to be some fixed value, e.g. 256px. One could always fill to a square image with black or gray or so, if one algorithm needs that.) However, also the downsampling itself can be done different ways. One can do simple averages over the original pixels that go into one downsampled one, one can use some weighting with window functions / kernels, one can Fourier transform the image and cut higher frequencies, etc.\nDoes anybody have a grasp on what the impact for modelling is?",
    "545904": "Well I am new to this Image recognition stuff ....Can anyone share a tutorial for this ?",
    "546004": "Hi Anish! Kudos to you for your great work in recognition. Although I didn't use your datasets last time since I was using Kernels solely and used the tar files in a download-train-delete kind of way. But here in this competition, the tar files are each of ~30 GB, which makes them impossible to use in kernels. If you happen to go forward and resize the train set, can you break them into tar files such that the uncompressed sizes of each of those are around 2-3 GB at max? In that way, a lot more people like me will be able to participate who otherwise have no such resources.",
    "546239": "Hey @mayukh18!\n\nFor sure. I've already shared the `test` and `validation` sets in a compressed format. I'll be re-sizing and then splitting up the `train` data-set into small zip files, and sharing them through a Kaggle dataset directly.\n\nI'll drop a message in this thread in the next couple of days when the download and re-size is complete on my end.",
    "550600": "You are amazing for doing this, thanks!",
    "552659": "&gt; Multiple people reached out and suggest I do proportional sizing though, since many models take square images as input, and it's easier for the end-users.\n\nI'd be interested to know more about the possibilities and the tradeoffs when resizing images and changing aspect ratios/filling up images. Could you kindly point to those comments you received and/or to some comprehensive resource on the topic? Many thanks in advance!",
    "555052": "anishagnihotri I am still not clear about the \"resizing\" thing. Could you point me to article/paper where the trade-off between size and quality is clearly explained.\n\nright now i am skimming through https://github.com/jarun/imgp",
    "555054": "just try to read through kaggle kernels and try googling/ stackoverflow for queries. There are tons of tutorials out there on computer vision."
  },
  "source": "meta"
}