{
  "id": 216263,
  "title": "Option to choose to exclude tfrecord or other formats",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/216263",
  "author_name": "ryches",
  "post_date": "2021-02-02T07:30:06.595000",
  "votes": 11,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I have noticed on several competitions now the data is provided in a few different formats, tfrecord + png/jpg/tif etc. It seems a bit wasteful to have all formats bundled together and give us no option to download specifically the files we want. Obviously, we could go to the menu and download what we want one by one but then we also have to organize directories the same way. </p>\n<p>This data is already fairly large but increasing it up to 160gb when I know I am going to delete the tfrecords after I spend the 6 hours (and I know in many parts of the world internet is even slower than mine)  it is going to take to download them is a bit inconvenient. I think it also likely scares off competitors who see such a high GB and think they wont be able to handle the requirements for the competition either because of compute or storage constraints. </p>\n<p>If an option to pick format could be provided that would be a big quality of life upgrade. </p>",
  "messages": [
    {
      "id": 1181881,
      "postDate": "2021-02-02T07:30:06.597Z",
      "content": "<p>I have noticed on several competitions now the data is provided in a few different formats, tfrecord + png/jpg/tif etc. It seems a bit wasteful to have all formats bundled together and give us no option to download specifically the files we want. Obviously, we could go to the menu and download what we want one by one but then we also have to organize directories the same way. </p>\n<p>This data is already fairly large but increasing it up to 160gb when I know I am going to delete the tfrecords after I spend the 6 hours (and I know in many parts of the world internet is even slower than mine)  it is going to take to download them is a bit inconvenient. I think it also likely scares off competitors who see such a high GB and think they wont be able to handle the requirements for the competition either because of compute or storage constraints. </p>\n<p>If an option to pick format could be provided that would be a big quality of life upgrade. </p>",
      "rawMarkdown": "I have noticed on several competitions now the data is provided in a few different formats, tfrecord + png/jpg/tif etc. It seems a bit wasteful to have all formats bundled together and give us no option to download specifically the files we want. Obviously, we could go to the menu and download what we want one by one but then we also have to organize directories the same way. \n\nThis data is already fairly large but increasing it up to 160gb when I know I am going to delete the tfrecords after I spend the 6 hours (and I know in many parts of the world internet is even slower than mine)  it is going to take to download them is a bit inconvenient. I think it also likely scares off competitors who see such a high GB and think they wont be able to handle the requirements for the competition either because of compute or storage constraints. \n\nIf an option to pick format could be provided that would be a big quality of life upgrade. ",
      "votes": 11
    },
    {
      "id": 1182610,
      "postDate": "2021-02-02T14:09:59.630Z",
      "content": "<p>Totally agree</p>\n<p>And using Kaggle API will download the whole, I cannot just  download a specific folder. </p>",
      "rawMarkdown": "Totally agree\n\nAnd using Kaggle API will download the whole, I cannot just  download a specific folder. ",
      "votes": 1
    },
    {
      "id": 1183684,
      "postDate": "2021-02-03T06:38:26.260Z",
      "content": "<p>Try this for downlaoding specific file.</p>\n<pre><code>kaggle competitions download -c human-protein-atlas-image-classification -f train.zip\nkaggle competitions download -c human-protein-atlas-image-classification -f test.zip\n\nmkdir -p dataset\n\nunzip train.zip -d dataset/train\nunzip test.zip -d dataset/test\n</code></pre>\n<p>kaggle competitions download -f</p>\n<p>-f FILE_NAME, --file FILE_NAME<br>\n {File name, all files downloaded if not provided}<br>\nsource <a href=\"https://github.com/Kaggle/kaggle-api\" target=\"_blank\">Kaggle api</a></p>",
      "rawMarkdown": "Try this for downlaoding specific file.\n\n```\nkaggle competitions download -c human-protein-atlas-image-classification -f train.zip\nkaggle competitions download -c human-protein-atlas-image-classification -f test.zip\n\nmkdir -p dataset\n\nunzip train.zip -d dataset/train\nunzip test.zip -d dataset/test\n```\n\n\nkaggle competitions download -f\n\n-f FILE_NAME, --file FILE_NAME\n {File name, all files downloaded if not provided}\nsource [Kaggle api](https://github.com/Kaggle/kaggle-api)"
    },
    {
      "id": 1183310,
      "postDate": "2021-02-02T22:20:40.110Z",
      "content": "<p>Unfortunately we don't have an easy way to fix this. We just asked the support team at Kaggle, and here is their response: </p>\n<blockquote>\n  <p>The TFRecords are included in the competition bundle because we've included TFRecords in the hidden dataset, so the competition mechanics require it. They are relatively small (~20GB) and shouldn't add significantly to the overall download size.</p>\n</blockquote>",
      "rawMarkdown": "Unfortunately we don't have an easy way to fix this. We just asked the support team at Kaggle, and here is their response: \n> The TFRecords are included in the competition bundle because we've included TFRecords in the hidden dataset, so the competition mechanics require it. They are relatively small (~20GB) and shouldn't add significantly to the overall download size.",
      "replies": [
        {
          "id": 1185292,
          "postDate": "2021-02-04T05:20:14.930Z",
          "content": "<p>My cynical translation: </p>\n<blockquote>\n  <p>The TFRecords are included in the competition bundle because we've included TFRecords in the hidden dataset, because we are now a subsidiary of Google, and Google TFRecords work really well on Google TPUs, and Google wants you to use TPUs.</p>\n</blockquote>\n<p>Added 5 Feb…<br>\nI think providing TFRecords (unbundled) to promote TPU use is a great idea. Just puzzled by the reply that someone gave to our competition host. </p>",
          "rawMarkdown": "My cynical translation: \n> The TFRecords are included in the competition bundle because we've included TFRecords in the hidden dataset, because we are now a subsidiary of Google, and Google TFRecords work really well on Google TPUs, and Google wants you to use TPUs.\n\nAdded 5 Feb...\nI think providing TFRecords (unbundled) to promote TPU use is a great idea. Just puzzled by the reply that someone gave to our competition host. ",
          "votes": 1
        },
        {
          "id": 1185322,
          "postDate": "2021-02-04T05:43:48.347Z",
          "content": "<p>I wasn't going to say it… But that's how I interpreted it as well. </p>",
          "rawMarkdown": "I wasn't going to say it... But that's how I interpreted it as well. "
        }
      ]
    },
    {
      "id": 1182077,
      "postDate": "2021-02-02T09:41:12.343Z",
      "content": "<p>I was under the impression that clicking the \"train\" folder followed by download would only download the train data without the corresponding tfrecords, or have I misunderstood?</p>",
      "rawMarkdown": "I was under the impression that clicking the \"train\" folder followed by download would only download the train data without the corresponding tfrecords, or have I misunderstood?",
      "replies": [
        {
          "id": 1182112,
          "postDate": "2021-02-02T10:01:45.130Z",
          "content": "<p>Yes, that is the case but then we would also need to click and download the test folder as well, and the train.csv and the sample_submission etc. And then organize the directory structure ourselves.</p>\n<p>It is a minor kaggle suggestion, not really specific to this competition just to give us a little more flexibility. </p>",
          "rawMarkdown": "Yes, that is the case but then we would also need to click and download the test folder as well, and the train.csv and the sample_submission etc. And then organize the directory structure ourselves.\n\nIt is a minor kaggle suggestion, not really specific to this competition just to give us a little more flexibility. ",
          "votes": 1
        },
        {
          "id": 1182139,
          "postDate": "2021-02-02T10:19:23.927Z",
          "content": "<p>I see what you mean. I can forward the feedback to Kaggle  staff at our next meeting.</p>",
          "rawMarkdown": "I see what you mean. I can forward the feedback to Kaggle  staff at our next meeting.",
          "votes": 2
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1182610,
      "author_name": "Serigne ",
      "author_url": "",
      "post_date": "2021-02-02T14:09:59.630000",
      "content": "<p>Totally agree</p>\n<p>And using Kaggle API will download the whole, I cannot just  download a specific folder. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1183684,
      "author_name": "Neo",
      "author_url": "",
      "post_date": "2021-02-03T06:38:26.260000",
      "content": "<p>Try this for downlaoding specific file.</p>\n<pre><code>kaggle competitions download -c human-protein-atlas-image-classification -f train.zip\nkaggle competitions download -c human-protein-atlas-image-classification -f test.zip\n\nmkdir -p dataset\n\nunzip train.zip -d dataset/train\nunzip test.zip -d dataset/test\n</code></pre>\n<p>kaggle competitions download -f</p>\n<p>-f FILE_NAME, --file FILE_NAME<br>\n {File name, all files downloaded if not provided}<br>\nsource <a href=\"https://github.com/Kaggle/kaggle-api\" target=\"_blank\">Kaggle api</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1183310,
      "author_name": "Wei Ouyang",
      "author_url": "",
      "post_date": "2021-02-02T22:20:40.110000",
      "content": "<p>Unfortunately we don't have an easy way to fix this. We just asked the support team at Kaggle, and here is their response: </p>\n<blockquote>\n  <p>The TFRecords are included in the competition bundle because we've included TFRecords in the hidden dataset, so the competition mechanics require it. They are relatively small (~20GB) and shouldn't add significantly to the overall download size.</p>\n</blockquote>",
      "votes": 0,
      "replies": [
        {
          "id": 1185292,
          "author_name": "JohnM",
          "author_url": "",
          "post_date": "2021-02-04T05:20:14.930000",
          "content": "<p>My cynical translation: </p>\n<blockquote>\n  <p>The TFRecords are included in the competition bundle because we've included TFRecords in the hidden dataset, because we are now a subsidiary of Google, and Google TFRecords work really well on Google TPUs, and Google wants you to use TPUs.</p>\n</blockquote>\n<p>Added 5 Feb…<br>\nI think providing TFRecords (unbundled) to promote TPU use is a great idea. Just puzzled by the reply that someone gave to our competition host. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1185322,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2021-02-04T05:43:48.347000",
          "content": "<p>I wasn't going to say it… But that's how I interpreted it as well. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1182077,
      "author_name": "Casper Winsnes",
      "author_url": "",
      "post_date": "2021-02-02T09:41:12.343000",
      "content": "<p>I was under the impression that clicking the \"train\" folder followed by download would only download the train data without the corresponding tfrecords, or have I misunderstood?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1182112,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2021-02-02T10:01:45.130000",
          "content": "<p>Yes, that is the case but then we would also need to click and download the test folder as well, and the train.csv and the sample_submission etc. And then organize the directory structure ourselves.</p>\n<p>It is a minor kaggle suggestion, not really specific to this competition just to give us a little more flexibility. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1182139,
          "author_name": "Casper Winsnes",
          "author_url": "",
          "post_date": "2021-02-02T10:19:23.927000",
          "content": "<p>I see what you mean. I can forward the feedback to Kaggle  staff at our next meeting.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1181881": "I have noticed on several competitions now the data is provided in a few different formats, tfrecord + png/jpg/tif etc. It seems a bit wasteful to have all formats bundled together and give us no option to download specifically the files we want. Obviously, we could go to the menu and download what we want one by one but then we also have to organize directories the same way. \n\nThis data is already fairly large but increasing it up to 160gb when I know I am going to delete the tfrecords after I spend the 6 hours (and I know in many parts of the world internet is even slower than mine)  it is going to take to download them is a bit inconvenient. I think it also likely scares off competitors who see such a high GB and think they wont be able to handle the requirements for the competition either because of compute or storage constraints. \n\nIf an option to pick format could be provided that would be a big quality of life upgrade. ",
    "1182610": "Totally agree\n\nAnd using Kaggle API will download the whole, I cannot just  download a specific folder. ",
    "1183684": "Try this for downlaoding specific file.\n\n```\nkaggle competitions download -c human-protein-atlas-image-classification -f train.zip\nkaggle competitions download -c human-protein-atlas-image-classification -f test.zip\n\nmkdir -p dataset\n\nunzip train.zip -d dataset/train\nunzip test.zip -d dataset/test\n```\n\n\nkaggle competitions download -f\n\n-f FILE_NAME, --file FILE_NAME\n {File name, all files downloaded if not provided}\nsource [Kaggle api](https://github.com/Kaggle/kaggle-api)",
    "1183310": "Unfortunately we don't have an easy way to fix this. We just asked the support team at Kaggle, and here is their response: \n> The TFRecords are included in the competition bundle because we've included TFRecords in the hidden dataset, so the competition mechanics require it. They are relatively small (~20GB) and shouldn't add significantly to the overall download size.",
    "1182077": "I was under the impression that clicking the \"train\" folder followed by download would only download the train data without the corresponding tfrecords, or have I misunderstood?"
  }
}