{
  "id": 207068,
  "title": "Not a feather file ",
  "url": "/competitions/riiid-test-answer-prediction/discussion/207068",
  "author_name": "",
  "post_date": "2020-12-28T04:51:21.127708900Z",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Why i get this error when i try to open feather file generated in gcp. <br>\nPandas 1.15 version.<br>\nIs it version issue. <br>\ni simply do df.to_feather( ) in gcp.</p>",
  "messages": [
    {
      "id": "1129100",
      "postDate": "12/28/2020 04:51:21",
      "content": "<p>Why i get this error when i try to open feather file generated in gcp. <br>\nPandas 1.15 version.<br>\nIs it version issue. <br>\ni simply do df.to_feather( ) in gcp.</p>",
      "rawMarkdown": "Why i get this error when i try to open feather file generated in gcp. \nPandas 1.15 version.\nIs it version issue. \ni simply do df.to_feather( ) in gcp.",
      "votes": null
    },
    {
      "id": "1129119",
      "postDate": "12/28/2020 05:11:17",
      "content": "<p>Probably. In the past, I've had issues arise with version conflicts between python and/or pandas. Same thing with pickles as well. Since the speed difference isn't significant—even when loading the 100M rows, I'd recommend saving yourself the headache (based on your posts, you've had some challenges) and just using parquet. They're both by Apache. Here <a href=\"https://stackoverflow.com/questions/48083405/what-are-the-differences-between-feather-and-parquet\" target=\"_blank\">is a list</a> of the differences by one of the devs.</p>",
      "rawMarkdown": "Probably. In the past, I've had issues arise with version conflicts between python and/or pandas. Same thing with pickles as well. Since the speed difference isn't significant—even when loading the 100M rows, I'd recommend saving yourself the headache (based on your posts, you've had some challenges) and just using parquet. They're both by Apache. Here [is a list](https://stackoverflow.com/questions/48083405/what-are-the-differences-between-feather-and-parquet) of the differences by one of the devs.",
      "votes": null
    },
    {
      "id": "1129134",
      "postDate": "12/28/2020 05:35:28",
      "content": "<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> thanks..<br>\nor probably i would just downgrade the version or another option u say is parquet.</p>\n<p>between any reason you are continuing to get good cv as i read in SAINT benchmark post  and but not yet got the lb accordingly ?</p>",
      "rawMarkdown": "authman thanks..\nor probably i would just downgrade the version or another option u say is parquet.\n\nbetween any reason you are continuing to get good cv as i read in SAINT benchmark post  and but not yet got the lb accordingly ?",
      "votes": null
    },
    {
      "id": "1129643",
      "postDate": "12/28/2020 13:42:33",
      "content": "<p>My FE pipeline is currently implemented 100% in pandas, which was a critical mistake on my part. It takes about 8 minutes to run on my machine for 80% of trainset, so I'd imagine taking triple that long on Kaggle. I haven't written my inference pipeline at all yet. If I could go back in time, I would have just used a single unified pipeline the way a lot of the published kernels are doing. I'll get around to it within the next day or two, in order to avoid a last minute rush like I experienced with DSB2019.</p>",
      "rawMarkdown": "My FE pipeline is currently implemented 100% in pandas, which was a critical mistake on my part. It takes about 8 minutes to run on my machine for 80% of trainset, so I'd imagine taking triple that long on Kaggle. I haven't written my inference pipeline at all yet. If I could go back in time, I would have just used a single unified pipeline the way a lot of the published kernels are doing. I'll get around to it within the next day or two, in order to avoid a last minute rush like I experienced with DSB2019.",
      "votes": null
    },
    {
      "id": "1130282",
      "postDate": "12/28/2020 22:53:13",
      "content": "<p>This is because the version of pyarrow in Kaggle notebook is much older (0.16.0) than the latest release (3.0.0).</p>\n<p>If you want to use feather-format, there are 2 options:</p>\n<ul>\n<li>Upgrading pyarrow in kaggle environment (<a href=\"https://www.kaggle.com/c/severstal-steel-defect-detection/discussion/113195\" target=\"_blank\">this</a> might help)</li>\n<li>Downgrading pyarrow of your GCP to 0.16.0, and re-create all feather files</li>\n</ul>\n<p>I think the second option is safer to avoid version conflict in kaggle environment.</p>",
      "rawMarkdown": "This is because the version of pyarrow in Kaggle notebook is much older (0.16.0) than the latest release (3.0.0).\n\nIf you want to use feather-format, there are 2 options:\n\n- Upgrading pyarrow in kaggle environment ([this](https://www.kaggle.com/c/severstal-steel-defect-detection/discussion/113195) might help)\n- Downgrading pyarrow of your GCP to 0.16.0, and re-create all feather files\n\nI think the second option is safer to avoid version conflict in kaggle environment.",
      "votes": null
    },
    {
      "id": "1130480",
      "postDate": "12/29/2020 04:57:45",
      "content": "<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a>  we can consider teamup. Our single SAINT score 78.6 , we are working on saint plus to make it score like saint ,78.8 is waht current plus is. If you think your saint plus can easily get you 79.2+ on lb  . then i think it can merge with our pipeline to give even better. if you think good team we can form. please post your lb with saint plus. <strong>Our target is nothing but Gold</strong>. We have guy now with a good lgb model. Lgb model is very necessary will always give  an increase of 0.007   on top of saint plus. </p>",
      "rawMarkdown": "authman  we can consider teamup. Our single SAINT score 78.6 , we are working on saint plus to make it score like saint ,78.8 is waht current plus is. If you think your saint plus can easily get you 79.2+ on lb  . then i think it can merge with our pipeline to give even better. if you think good team we can form. please post your lb with saint plus. **Our target is nothing but Gold**. We have guy now with a good lgb model. Lgb model is very necessary will always give  an increase of 0.007   on top of saint plus.",
      "votes": null
    },
    {
      "id": "1165392",
      "postDate": "01/22/2021 23:53:02",
      "content": "<p>RIIID is history, but just in case anyone sees this thread again, you should try:</p>\n<p>pd.to_feather(path, version=1)</p>\n<p>Pandas defaults to saving with version=2, which cannot be read using libraries installed on kaggle. </p>",
      "rawMarkdown": "RIIID is history, but just in case anyone sees this thread again, you should try:\n\npd.to_feather(path, version=1)\n\nPandas defaults to saving with version=2, which cannot be read using libraries installed on kaggle.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1129119,
      "author_name": "authman",
      "author_url": "",
      "post_date": "12/28/2020 05:11:17",
      "content": "<p>Probably. In the past, I've had issues arise with version conflicts between python and/or pandas. Same thing with pickles as well. Since the speed difference isn't significant—even when loading the 100M rows, I'd recommend saving yourself the headache (based on your posts, you've had some challenges) and just using parquet. They're both by Apache. Here <a href=\"https://stackoverflow.com/questions/48083405/what-are-the-differences-between-feather-and-parquet\" target=\"_blank\">is a list</a> of the differences by one of the devs.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1129134,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/28/2020 05:35:28",
          "content": "<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> thanks..<br>\nor probably i would just downgrade the version or another option u say is parquet.</p>\n<p>between any reason you are continuing to get good cv as i read in SAINT benchmark post  and but not yet got the lb accordingly ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1129643,
          "author_name": "authman",
          "author_url": "",
          "post_date": "12/28/2020 13:42:33",
          "content": "<p>My FE pipeline is currently implemented 100% in pandas, which was a critical mistake on my part. It takes about 8 minutes to run on my machine for 80% of trainset, so I'd imagine taking triple that long on Kaggle. I haven't written my inference pipeline at all yet. If I could go back in time, I would have just used a single unified pipeline the way a lot of the published kernels are doing. I'll get around to it within the next day or two, in order to avoid a last minute rush like I experienced with DSB2019.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1130480,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/29/2020 04:57:45",
          "content": "<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a>  we can consider teamup. Our single SAINT score 78.6 , we are working on saint plus to make it score like saint ,78.8 is waht current plus is. If you think your saint plus can easily get you 79.2+ on lb  . then i think it can merge with our pipeline to give even better. if you think good team we can form. please post your lb with saint plus. <strong>Our target is nothing but Gold</strong>. We have guy now with a good lgb model. Lgb model is very necessary will always give  an increase of 0.007   on top of saint plus. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1130282,
      "author_name": "nyanpn",
      "author_url": "",
      "post_date": "12/28/2020 22:53:13",
      "content": "<p>This is because the version of pyarrow in Kaggle notebook is much older (0.16.0) than the latest release (3.0.0).</p>\n<p>If you want to use feather-format, there are 2 options:</p>\n<ul>\n<li>Upgrading pyarrow in kaggle environment (<a href=\"https://www.kaggle.com/c/severstal-steel-defect-detection/discussion/113195\" target=\"_blank\">this</a> might help)</li>\n<li>Downgrading pyarrow of your GCP to 0.16.0, and re-create all feather files</li>\n</ul>\n<p>I think the second option is safer to avoid version conflict in kaggle environment.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1165392,
      "author_name": "jakethomsen",
      "author_url": "",
      "post_date": "01/22/2021 23:53:02",
      "content": "<p>RIIID is history, but just in case anyone sees this thread again, you should try:</p>\n<p>pd.to_feather(path, version=1)</p>\n<p>Pandas defaults to saving with version=2, which cannot be read using libraries installed on kaggle. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1129100": "Why i get this error when i try to open feather file generated in gcp. \nPandas 1.15 version.\nIs it version issue. \ni simply do df.to_feather( ) in gcp.",
    "1129119": "Probably. In the past, I've had issues arise with version conflicts between python and/or pandas. Same thing with pickles as well. Since the speed difference isn't significant—even when loading the 100M rows, I'd recommend saving yourself the headache (based on your posts, you've had some challenges) and just using parquet. They're both by Apache. Here [is a list](https://stackoverflow.com/questions/48083405/what-are-the-differences-between-feather-and-parquet) of the differences by one of the devs.",
    "1129134": "authman thanks..\nor probably i would just downgrade the version or another option u say is parquet.\n\nbetween any reason you are continuing to get good cv as i read in SAINT benchmark post  and but not yet got the lb accordingly ?",
    "1129643": "My FE pipeline is currently implemented 100% in pandas, which was a critical mistake on my part. It takes about 8 minutes to run on my machine for 80% of trainset, so I'd imagine taking triple that long on Kaggle. I haven't written my inference pipeline at all yet. If I could go back in time, I would have just used a single unified pipeline the way a lot of the published kernels are doing. I'll get around to it within the next day or two, in order to avoid a last minute rush like I experienced with DSB2019.",
    "1130282": "This is because the version of pyarrow in Kaggle notebook is much older (0.16.0) than the latest release (3.0.0).\n\nIf you want to use feather-format, there are 2 options:\n\n- Upgrading pyarrow in kaggle environment ([this](https://www.kaggle.com/c/severstal-steel-defect-detection/discussion/113195) might help)\n- Downgrading pyarrow of your GCP to 0.16.0, and re-create all feather files\n\nI think the second option is safer to avoid version conflict in kaggle environment.",
    "1130480": "authman  we can consider teamup. Our single SAINT score 78.6 , we are working on saint plus to make it score like saint ,78.8 is waht current plus is. If you think your saint plus can easily get you 79.2+ on lb  . then i think it can merge with our pipeline to give even better. if you think good team we can form. please post your lb with saint plus. **Our target is nothing but Gold**. We have guy now with a good lgb model. Lgb model is very necessary will always give  an increase of 0.007   on top of saint plus.",
    "1165392": "RIIID is history, but just in case anyone sees this thread again, you should try:\n\npd.to_feather(path, version=1)\n\nPandas defaults to saving with version=2, which cannot be read using libraries installed on kaggle."
  },
  "source": "meta"
}