{
  "id": 94650,
  "title": "[0.3-3GB] Re-sized and compressed test data-set (512x, 256x, 128x)",
  "url": "/competitions/open-images-2019-object-detection/discussion/94650",
  "author_name": "Anish Agnihotri",
  "post_date": "2019-06-05T23:52:01.307000",
  "votes": 33,
  "comment_count": 11,
  "views": 0,
  "content": "<p><strong>Why download 9.7GB when you can download ~400MB?</strong> <em>Hooray</em></p>\n\n<p>Hey everyone,</p>\n\n<p>It looks like this is another one of those competitions with giant data-sets. In order to make it easier for everyone to compete, I am going to be re-sizing the <code>train</code>, <code>validation</code>, and <code>testing</code> data for the competition (for people who have limited download bandwidth or resources. I'm starting by sharing various sizes of the <code>testing</code> data-set, with more coming soon.</p>\n\n<p><strong>Note:</strong> These dimensions are not proportional, since the images are not exact squares. If there is demand for a forced-proportional data-set, I'd be happy to share that as well.</p>\n\n<h1>Downloads</h1>\n\n<p><strong>Original data-set</strong>\nUncompressed size: <code>~20GB</code></p>\n\n<p><strong><code>512x</code> data-set</strong>\nUncompressed size: <code>3.4GB</code>\nDownload link: <a href=\"https://drive.google.com/file/d/15AATyBwhP45GXd2CquVtFoNgsRldMG9z/view?usp=sharing\">Drive</a></p>\n\n<p><strong><code>256x</code> data-set</strong>\nUncompressed size: <code>1.2GB</code>\nDownload link: <a href=\"https://drive.google.com/file/d/1vlxh1mqmXL3GT_FQ3FwG9WpeHAfTF62u/view?usp=sharing\">Drive</a></p>\n\n<p><strong><code>128x</code> data-set</strong>\nUncompressed size: <code>~480MB</code>\nDownload link: <a href=\"https://drive.google.com/file/d/1LTkg1OLORQx7rKm_X-Etz-Ck-ayiKSD6/view?usp=sharing\">Drive</a></p>\n\n<h1>Additional information</h1>\n\n<p>For <a href=\"https://www.kaggle.com/c/landmark-recognition-2019/discussion/91770\">process, limitations, and specific download instructions</a> (if you're using a CLI, these might be easier), please refer to the post I made for the last competition of this sort. The instructions are the same.</p>\n\n<p>Cheers!</p>\n\n<p>Anish</p>",
  "messages": [
    {
      "id": 544824,
      "postDate": "2019-06-05T23:52:01.307Z",
      "content": "<p><strong>Why download 9.7GB when you can download ~400MB?</strong> <em>Hooray</em></p>\n\n<p>Hey everyone,</p>\n\n<p>It looks like this is another one of those competitions with giant data-sets. In order to make it easier for everyone to compete, I am going to be re-sizing the <code>train</code>, <code>validation</code>, and <code>testing</code> data for the competition (for people who have limited download bandwidth or resources. I'm starting by sharing various sizes of the <code>testing</code> data-set, with more coming soon.</p>\n\n<p><strong>Note:</strong> These dimensions are not proportional, since the images are not exact squares. If there is demand for a forced-proportional data-set, I'd be happy to share that as well.</p>\n\n<h1>Downloads</h1>\n\n<p><strong>Original data-set</strong>\nUncompressed size: <code>~20GB</code></p>\n\n<p><strong><code>512x</code> data-set</strong>\nUncompressed size: <code>3.4GB</code>\nDownload link: <a href=\"https://drive.google.com/file/d/15AATyBwhP45GXd2CquVtFoNgsRldMG9z/view?usp=sharing\">Drive</a></p>\n\n<p><strong><code>256x</code> data-set</strong>\nUncompressed size: <code>1.2GB</code>\nDownload link: <a href=\"https://drive.google.com/file/d/1vlxh1mqmXL3GT_FQ3FwG9WpeHAfTF62u/view?usp=sharing\">Drive</a></p>\n\n<p><strong><code>128x</code> data-set</strong>\nUncompressed size: <code>~480MB</code>\nDownload link: <a href=\"https://drive.google.com/file/d/1LTkg1OLORQx7rKm_X-Etz-Ck-ayiKSD6/view?usp=sharing\">Drive</a></p>\n\n<h1>Additional information</h1>\n\n<p>For <a href=\"https://www.kaggle.com/c/landmark-recognition-2019/discussion/91770\">process, limitations, and specific download instructions</a> (if you're using a CLI, these might be easier), please refer to the post I made for the last competition of this sort. The instructions are the same.</p>\n\n<p>Cheers!</p>\n\n<p>Anish</p>",
      "rawMarkdown": "**Why download 9.7GB when you can download ~400MB?** *Hooray*\n\nHey everyone,\n\nIt looks like this is another one of those competitions with giant data-sets. In order to make it easier for everyone to compete, I am going to be re-sizing the `train`, `validation`, and `testing` data for the competition (for people who have limited download bandwidth or resources. I'm starting by sharing various sizes of the `testing` data-set, with more coming soon.\n\n**Note:** These dimensions are not proportional, since the images are not exact squares. If there is demand for a forced-proportional data-set, I'd be happy to share that as well.\n\n# Downloads\n\n**Original data-set**\nUncompressed size: `~20GB`\n\n**`512x` data-set**\nUncompressed size: `3.4GB`\nDownload link: [Drive](https://drive.google.com/file/d/15AATyBwhP45GXd2CquVtFoNgsRldMG9z/view?usp=sharing)\n\n**`256x` data-set**\nUncompressed size: `1.2GB`\nDownload link: [Drive](https://drive.google.com/file/d/1vlxh1mqmXL3GT_FQ3FwG9WpeHAfTF62u/view?usp=sharing)\n\n**`128x` data-set**\nUncompressed size: `~480MB`\nDownload link: [Drive](https://drive.google.com/file/d/1LTkg1OLORQx7rKm_X-Etz-Ck-ayiKSD6/view?usp=sharing)\n\n# Additional information\nFor [process, limitations, and specific download instructions](https://www.kaggle.com/c/landmark-recognition-2019/discussion/91770) (if you're using a CLI, these might be easier), please refer to the post I made for the last competition of this sort. The instructions are the same.\n\nCheers!\n\nAnish\n\n",
      "votes": 31
    },
    {
      "id": 554940,
      "postDate": "2019-06-18T07:54:06.340Z",
      "content": "<p>TKS. Just what I want. Beside of GoogleDriver, any other way to get the data?</p>",
      "rawMarkdown": "TKS. Just what I want. Beside of GoogleDriver, any other way to get the data?"
    },
    {
      "id": 548090,
      "postDate": "2019-06-08T18:56:27.113Z",
      "content": "<p>When will the train data be uploaded ?</p>",
      "rawMarkdown": "When will the train data be uploaded ?"
    },
    {
      "id": 547402,
      "postDate": "2019-06-07T17:00:22.527Z",
      "content": "<p>Thank you very much! Since I have limited resources,  that would help me a lot! :)</p>",
      "rawMarkdown": "Thank you very much! Since I have limited resources,  that would help me a lot! :)"
    },
    {
      "id": 546299,
      "postDate": "2019-06-06T12:53:05.093Z",
      "content": "<p>scalling down the images will loss the quality, so it will become almost impossible to capture small detail or to detect small objects.\nAm i right?</p>",
      "rawMarkdown": "scalling down the images will loss the quality, so it will become almost impossible to capture small detail or to detect small objects.\nAm i right?",
      "replies": [
        {
          "id": 546331,
          "postDate": "2019-06-06T13:31:36.083Z",
          "content": "<p>That's correct. While you shouldn't lose much quality on the <code>512px</code> re-sizes, you will be losing a significant amount of data going to any sizes smaller than that.</p>\n\n<p>This re-size is simply for individuals who cannot participate otherwise (since the data-set is over 600GB), want to do EDA, or need to use a Kernel for their processes (since data-sets have a 20GB limit). It's still better than nothing for these individuals.</p>",
          "rawMarkdown": "That's correct. While you shouldn't lose much quality on the `512px` re-sizes, you will be losing a significant amount of data going to any sizes smaller than that.\n\nThis re-size is simply for individuals who cannot participate otherwise (since the data-set is over 600GB), want to do EDA, or need to use a Kernel for their processes (since data-sets have a 20GB limit). It's still better than nothing for these individuals.",
          "votes": 2
        }
      ]
    },
    {
      "id": 545914,
      "postDate": "2019-06-06T03:27:47.143Z",
      "content": "<p>Excellent!</p>",
      "rawMarkdown": "Excellent!",
      "replies": [
        {
          "id": 545933,
          "postDate": "2019-06-06T03:58:30.097Z",
          "content": "<p>Thanks! Glad I could help. I'll be uploading the <code>train</code> data in the next couple of days as well.</p>",
          "rawMarkdown": "Thanks! Glad I could help. I'll be uploading the `train` data in the next couple of days as well.",
          "votes": 1
        }
      ]
    },
    {
      "id": 555950,
      "postDate": "2019-06-19T16:20:13.487Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 552736,
      "postDate": "2019-06-14T12:43:12.467Z",
      "content": "<p>Thanks a lot</p>",
      "rawMarkdown": "Thanks a lot"
    },
    {
      "id": 551056,
      "postDate": "2019-06-12T09:41:05.353Z",
      "content": "<p>Thank you very much!</p>",
      "rawMarkdown": "Thank you very much!"
    },
    {
      "id": 547725,
      "postDate": "2019-06-08T07:25:40.990Z",
      "content": "<p>Thank you . </p>",
      "rawMarkdown": "Thank you . "
    }
  ],
  "comments": [
    {
      "id": 554940,
      "author_name": "Arthur Zhou",
      "author_url": "",
      "post_date": "2019-06-18T07:54:06.340000",
      "content": "<p>TKS. Just what I want. Beside of GoogleDriver, any other way to get the data?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 548090,
      "author_name": "Mohamed Ramzy",
      "author_url": "",
      "post_date": "2019-06-08T18:56:27.113000",
      "content": "<p>When will the train data be uploaded ?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 547402,
      "author_name": "Tiago de Miranda Leite",
      "author_url": "",
      "post_date": "2019-06-07T17:00:22.527000",
      "content": "<p>Thank you very much! Since I have limited resources,  that would help me a lot! :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 546299,
      "author_name": "faheem",
      "author_url": "",
      "post_date": "2019-06-06T12:53:05.093000",
      "content": "<p>scalling down the images will loss the quality, so it will become almost impossible to capture small detail or to detect small objects.\nAm i right?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 546331,
          "author_name": "Anish Agnihotri",
          "author_url": "",
          "post_date": "2019-06-06T13:31:36.083000",
          "content": "<p>That's correct. While you shouldn't lose much quality on the <code>512px</code> re-sizes, you will be losing a significant amount of data going to any sizes smaller than that.</p>\n\n<p>This re-size is simply for individuals who cannot participate otherwise (since the data-set is over 600GB), want to do EDA, or need to use a Kernel for their processes (since data-sets have a 20GB limit). It's still better than nothing for these individuals.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 545914,
      "author_name": "interneuron",
      "author_url": "",
      "post_date": "2019-06-06T03:27:47.143000",
      "content": "<p>Excellent!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 545933,
          "author_name": "Anish Agnihotri",
          "author_url": "",
          "post_date": "2019-06-06T03:58:30.097000",
          "content": "<p>Thanks! Glad I could help. I'll be uploading the <code>train</code> data in the next couple of days as well.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 555950,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-19T16:20:13.487000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 552736,
      "author_name": "zxyu",
      "author_url": "",
      "post_date": "2019-06-14T12:43:12.467000",
      "content": "<p>Thanks a lot</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 551056,
      "author_name": "Shiva Manhar",
      "author_url": "",
      "post_date": "2019-06-12T09:41:05.353000",
      "content": "<p>Thank you very much!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 547725,
      "author_name": "PratikChatse",
      "author_url": "",
      "post_date": "2019-06-08T07:25:40.990000",
      "content": "<p>Thank you . </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "544824": "**Why download 9.7GB when you can download ~400MB?** *Hooray*\n\nHey everyone,\n\nIt looks like this is another one of those competitions with giant data-sets. In order to make it easier for everyone to compete, I am going to be re-sizing the `train`, `validation`, and `testing` data for the competition (for people who have limited download bandwidth or resources. I'm starting by sharing various sizes of the `testing` data-set, with more coming soon.\n\n**Note:** These dimensions are not proportional, since the images are not exact squares. If there is demand for a forced-proportional data-set, I'd be happy to share that as well.\n\n# Downloads\n\n**Original data-set**\nUncompressed size: `~20GB`\n\n**`512x` data-set**\nUncompressed size: `3.4GB`\nDownload link: [Drive](https://drive.google.com/file/d/15AATyBwhP45GXd2CquVtFoNgsRldMG9z/view?usp=sharing)\n\n**`256x` data-set**\nUncompressed size: `1.2GB`\nDownload link: [Drive](https://drive.google.com/file/d/1vlxh1mqmXL3GT_FQ3FwG9WpeHAfTF62u/view?usp=sharing)\n\n**`128x` data-set**\nUncompressed size: `~480MB`\nDownload link: [Drive](https://drive.google.com/file/d/1LTkg1OLORQx7rKm_X-Etz-Ck-ayiKSD6/view?usp=sharing)\n\n# Additional information\nFor [process, limitations, and specific download instructions](https://www.kaggle.com/c/landmark-recognition-2019/discussion/91770) (if you're using a CLI, these might be easier), please refer to the post I made for the last competition of this sort. The instructions are the same.\n\nCheers!\n\nAnish\n\n",
    "554940": "TKS. Just what I want. Beside of GoogleDriver, any other way to get the data?",
    "548090": "When will the train data be uploaded ?",
    "547402": "Thank you very much! Since I have limited resources,  that would help me a lot! :)",
    "546299": "scalling down the images will loss the quality, so it will become almost impossible to capture small detail or to detect small objects.\nAm i right?",
    "545914": "Excellent!",
    "555950": "",
    "552736": "Thanks a lot",
    "551056": "Thank you very much!",
    "547725": "Thank you . "
  }
}