{
  "id": 466975,
  "title": "Broken Test Parquet Files and will it be fixed during the competition window?",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/466975",
  "author_name": "",
  "post_date": "2024-01-10T15:56:50.494659900Z",
  "votes": 5,
  "comment_count": 1,
  "views": 0,
  "content": "<p>The test EEG file \"3911565283.parquet\" is not actually a parquet file and must be read using pd.read_csv() instead of pd.read_parquet().  While it is not an issue to use read_csv instead, do we know if this will be fixed at some point?  I ask because I am worried that some of the hidden test files in this competition may actually be in parquet format while \"3911565283.parquet\" is a CSV and will then cause errors when submitting.  It is also easier to have the training and testing data in the same format for the sake of consistency.</p>",
  "messages": [
    {
      "id": "2595733",
      "postDate": "01/10/2024 15:56:50",
      "content": "<p>The test EEG file \"3911565283.parquet\" is not actually a parquet file and must be read using pd.read_csv() instead of pd.read_parquet().  While it is not an issue to use read_csv instead, do we know if this will be fixed at some point?  I ask because I am worried that some of the hidden test files in this competition may actually be in parquet format while \"3911565283.parquet\" is a CSV and will then cause errors when submitting.  It is also easier to have the training and testing data in the same format for the sake of consistency.</p>",
      "rawMarkdown": "The test EEG file \"3911565283.parquet\" is not actually a parquet file and must be read using pd.read_csv() instead of pd.read_parquet().  While it is not an issue to use read_csv instead, do we know if this will be fixed at some point?  I ask because I am worried that some of the hidden test files in this competition may actually be in parquet format while \"3911565283.parquet\" is a CSV and will then cause errors when submitting.  It is also easier to have the training and testing data in the same format for the sake of consistency.",
      "votes": null
    },
    {
      "id": "2595979",
      "postDate": "01/10/2024 18:59:22",
      "content": "<p>Thanks for flagging this. I've posted a patch.</p>",
      "rawMarkdown": "Thanks for flagging this. I've posted a patch.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2595979,
      "author_name": "sohier",
      "author_url": "",
      "post_date": "01/10/2024 18:59:22",
      "content": "<p>Thanks for flagging this. I've posted a patch.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2595733": "The test EEG file \"3911565283.parquet\" is not actually a parquet file and must be read using pd.read_csv() instead of pd.read_parquet().  While it is not an issue to use read_csv instead, do we know if this will be fixed at some point?  I ask because I am worried that some of the hidden test files in this competition may actually be in parquet format while \"3911565283.parquet\" is a CSV and will then cause errors when submitting.  It is also easier to have the training and testing data in the same format for the sake of consistency.",
    "2595979": "Thanks for flagging this. I've posted a patch."
  },
  "source": "meta"
}