{
  "id": 80290,
  "title": "what use is the train and test parquet?",
  "url": "/competitions/vsb-power-line-fault-detection/discussion/80290",
  "author_name": "Richie",
  "post_date": "2019-02-12T11:10:58.055000",
  "votes": -3,
  "comment_count": 6,
  "views": 0,
  "content": "<p>what are we to do with the train and test parquet?  from the overview and data statement it seems the train and test .csv files are sufficient.</p>",
  "messages": [
    {
      "id": 470241,
      "postDate": "2019-02-12T16:33:21.110Z",
      "content": "<p>signal data which you need to use in the training phase and prediction phase is in parquet.\n.csv files only contain meta data (id_measurement, signal_id, phase, target).</p>\n\n<p>you can check descriptions about data in <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/data\">https://www.kaggle.com/c/vsb-power-line-fault-detection/data</a>. </p>",
      "rawMarkdown": "signal data which you need to use in the training phase and prediction phase is in parquet.\n.csv files only contain meta data (id_measurement, signal_id, phase, target).\n\nyou can check descriptions about data in https://www.kaggle.com/c/vsb-power-line-fault-detection/data. ",
      "votes": 1
    },
    {
      "id": 470128,
      "postDate": "2019-02-12T12:39:36.277Z",
      "content": "<p>The data is in the parquet, the metadata and labels in the csv.</p>",
      "rawMarkdown": "The data is in the parquet, the metadata and labels in the csv.",
      "votes": 1,
      "replies": [
        {
          "id": 474287,
          "postDate": "2019-02-19T07:04:48.893Z",
          "content": "<p>can i import parquets into R, I want to use R and have only used csv, sql and excel datasets.  I also observed the size of the parquet files are about 10gb.</p>\n\n<p>I get it now, it seems I have to understand SparkR to help me import and work with these datasets.</p>",
          "rawMarkdown": "can i import parquets into R, I want to use R and have only used csv, sql and excel datasets.  I also observed the size of the parquet files are about 10gb.\n\nI get it now, it seems I have to understand SparkR to help me import and work with these datasets.\n\n"
        },
        {
          "id": 474823,
          "postDate": "2019-02-19T21:07:25.323Z",
          "content": "<p>I am also working in R, although I use the R \"reticulate\" package to get to Python 3 from R if/when I need to.  I tried using Spark from R to read the parquet files but gave up due to too many problems.  What finally worked for me was to read the parquet files into Python and write them out as feather files.  This only needed to be done once and required the Python packages pyarrow.parquet, pandas, and feather.  Once I had the feather files, I could read them into R using the read_feather function from the R feather package.   read_feather is quite efficient and can read an entire feather file or whatever columns you desire.</p>\n\n<p>Hope this helps.</p>",
          "rawMarkdown": "I am also working in R, although I use the R \"reticulate\" package to get to Python 3 from R if/when I need to.  I tried using Spark from R to read the parquet files but gave up due to too many problems.  What finally worked for me was to read the parquet files into Python and write them out as feather files.  This only needed to be done once and required the Python packages pyarrow.parquet, pandas, and feather.  Once I had the feather files, I could read them into R using the read_feather function from the R feather package.   read_feather is quite efficient and can read an entire feather file or whatever columns you desire.\n\nHope this helps.\n\n\n\n      \n\n"
        },
        {
          "id": 475158,
          "postDate": "2019-02-20T11:05:12.743Z",
          "content": "<p>from your response seems it would be easier doing this in python then.  </p>",
          "rawMarkdown": "from your response seems it would be easier doing this in python then.  "
        }
      ]
    },
    {
      "id": 470093,
      "postDate": "2019-02-12T11:10:58.057Z",
      "content": "<p>what are we to do with the train and test parquet?  from the overview and data statement it seems the train and test .csv files are sufficient.</p>",
      "rawMarkdown": "what are we to do with the train and test parquet?  from the overview and data statement it seems the train and test .csv files are sufficient.",
      "votes": -3
    },
    {
      "id": 470126,
      "postDate": "2019-02-12T12:38:33.257Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 470241,
      "author_name": "Kim",
      "author_url": "",
      "post_date": "2019-02-12T16:33:21.110000",
      "content": "<p>signal data which you need to use in the training phase and prediction phase is in parquet.\n.csv files only contain meta data (id_measurement, signal_id, phase, target).</p>\n\n<p>you can check descriptions about data in <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/data\">https://www.kaggle.com/c/vsb-power-line-fault-detection/data</a>. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 470128,
      "author_name": "Joop",
      "author_url": "",
      "post_date": "2019-02-12T12:39:36.277000",
      "content": "<p>The data is in the parquet, the metadata and labels in the csv.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 474287,
          "author_name": "Richie",
          "author_url": "",
          "post_date": "2019-02-19T07:04:48.893000",
          "content": "<p>can i import parquets into R, I want to use R and have only used csv, sql and excel datasets.  I also observed the size of the parquet files are about 10gb.</p>\n\n<p>I get it now, it seems I have to understand SparkR to help me import and work with these datasets.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 474823,
          "author_name": "David J. Slate",
          "author_url": "",
          "post_date": "2019-02-19T21:07:25.323000",
          "content": "<p>I am also working in R, although I use the R \"reticulate\" package to get to Python 3 from R if/when I need to.  I tried using Spark from R to read the parquet files but gave up due to too many problems.  What finally worked for me was to read the parquet files into Python and write them out as feather files.  This only needed to be done once and required the Python packages pyarrow.parquet, pandas, and feather.  Once I had the feather files, I could read them into R using the read_feather function from the R feather package.   read_feather is quite efficient and can read an entire feather file or whatever columns you desire.</p>\n\n<p>Hope this helps.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 475158,
          "author_name": "Richie",
          "author_url": "",
          "post_date": "2019-02-20T11:05:12.743000",
          "content": "<p>from your response seems it would be easier doing this in python then.  </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 470126,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-12T12:38:33.257000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "470241": "signal data which you need to use in the training phase and prediction phase is in parquet.\n.csv files only contain meta data (id_measurement, signal_id, phase, target).\n\nyou can check descriptions about data in https://www.kaggle.com/c/vsb-power-line-fault-detection/data. ",
    "470128": "The data is in the parquet, the metadata and labels in the csv.",
    "470093": "what are we to do with the train and test parquet?  from the overview and data statement it seems the train and test .csv files are sufficient.",
    "470126": ""
  }
}