{
  "id": 122716,
  "title": "Why was parquet file format chosen for this competition?",
  "url": "/competitions/bengaliai-cv19/discussion/122716",
  "author_name": "",
  "post_date": "2019-12-22T13:02:28.643889600Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "",
  "messages": [
    {
      "id": "700691",
      "postDate": "12/22/2019 13:02:28",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "700860",
      "postDate": "12/22/2019 18:29:20",
      "content": "<p>Probably due to size and the necessary compression.</p>",
      "rawMarkdown": "Probably due to size and the necessary compression.",
      "votes": null
    },
    {
      "id": "707452",
      "postDate": "12/31/2019 21:25:16",
      "content": "<p>I am not sure it's the reason: saving each image individually in a png file gives a total size of 1.8 Go (compressing to tar.gz does not reduce more), so much less than the 5.1 Go of the .parquet files.</p>",
      "rawMarkdown": "I am not sure it's the reason: saving each image individually in a png file gives a total size of 1.8 Go (compressing to tar.gz does not reduce more), so much less than the 5.1 Go of the .parquet files.",
      "votes": null
    },
    {
      "id": "707604",
      "postDate": "01/01/2020 06:10:31",
      "content": "<p>My preprocessed image dataset is only ~0.5 GB. I guess parquet is chosen because this format is used by organizers in their pipline.</p>",
      "rawMarkdown": "My preprocessed image dataset is only ~0.5 GB. I guess parquet is chosen because this format is used by organizers in their pipline.",
      "votes": null
    },
    {
      "id": "721672",
      "postDate": "01/17/2020 15:40:10",
      "content": "<p>Hi <a href=\"/iafoss\">@iafoss</a>, I think organizers should have provided a more standard format. I only could make some tests and publish a related <a href=\"https://www.kaggle.com/jmartindelasierra/class-activation-mapping-in-handwritten-bengali\">notebook</a> thanks to your preprocessed dataset but I can't make any progress and a submission since I'm unable to open this files in R.</p>",
      "rawMarkdown": "Hi @iafoss, I think organizers should have provided a more standard format. I only could make some tests and publish a related [notebook](https://www.kaggle.com/jmartindelasierra/class-activation-mapping-in-handwritten-bengali) thanks to your preprocessed dataset but I can't make any progress and a submission since I'm unable to open this files in R.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 700860,
      "author_name": "konradb",
      "author_url": "",
      "post_date": "12/22/2019 18:29:20",
      "content": "<p>Probably due to size and the necessary compression.</p>",
      "votes": null,
      "replies": [
        {
          "id": 707452,
          "author_name": "melsophos",
          "author_url": "",
          "post_date": "12/31/2019 21:25:16",
          "content": "<p>I am not sure it's the reason: saving each image individually in a png file gives a total size of 1.8 Go (compressing to tar.gz does not reduce more), so much less than the 5.1 Go of the .parquet files.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 707604,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "01/01/2020 06:10:31",
          "content": "<p>My preprocessed image dataset is only ~0.5 GB. I guess parquet is chosen because this format is used by organizers in their pipline.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 721672,
          "author_name": "jmartindelasierra",
          "author_url": "",
          "post_date": "01/17/2020 15:40:10",
          "content": "<p>Hi <a href=\"/iafoss\">@iafoss</a>, I think organizers should have provided a more standard format. I only could make some tests and publish a related <a href=\"https://www.kaggle.com/jmartindelasierra/class-activation-mapping-in-handwritten-bengali\">notebook</a> thanks to your preprocessed dataset but I can't make any progress and a submission since I'm unable to open this files in R.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "700691": "",
    "700860": "Probably due to size and the necessary compression.",
    "707452": "I am not sure it's the reason: saving each image individually in a png file gives a total size of 1.8 Go (compressing to tar.gz does not reduce more), so much less than the 5.1 Go of the .parquet files.",
    "707604": "My preprocessed image dataset is only ~0.5 GB. I guess parquet is chosen because this format is used by organizers in their pipline.",
    "721672": "Hi @iafoss, I think organizers should have provided a more standard format. I only could make some tests and publish a related [notebook](https://www.kaggle.com/jmartindelasierra/class-activation-mapping-in-handwritten-bengali) thanks to your preprocessed dataset but I can't make any progress and a submission since I'm unable to open this files in R."
  },
  "source": "meta"
}