{
  "id": 55581,
  "title": "saving/loading feature from local disk",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/55581",
  "author_name": "Cheng",
  "post_date": "2018-04-28T21:19:43.009000",
  "votes": 5,
  "comment_count": 19,
  "views": 0,
  "content": "<p>I am experiencing a very strange bug: when I save the generated feature to local disk as in</p>\n\n<p><a href=\"https://www.kaggle.com/bk0000/non-blending-lightgbm-model-lb-0-977\">https://www.kaggle.com/bk0000/non-blending-lightgbm-model-lb-0-977</a></p>\n\n<p>and load back, my performance would drop significantly (&gt;0.02 on pubic LB), while the validation performance remains close, even if I run the identical model. Does anyone experiencing similar issue? Thank you!</p>",
  "messages": [
    {
      "id": 320510,
      "postDate": "2018-04-28T21:19:43.010Z",
      "content": "<p>I am experiencing a very strange bug: when I save the generated feature to local disk as in</p>\n\n<p><a href=\"https://www.kaggle.com/bk0000/non-blending-lightgbm-model-lb-0-977\">https://www.kaggle.com/bk0000/non-blending-lightgbm-model-lb-0-977</a></p>\n\n<p>and load back, my performance would drop significantly (&gt;0.02 on pubic LB), while the validation performance remains close, even if I run the identical model. Does anyone experiencing similar issue? Thank you!</p>",
      "rawMarkdown": "I am experiencing a very strange bug: when I save the generated feature to local disk as in\n\nhttps://www.kaggle.com/bk0000/non-blending-lightgbm-model-lb-0-977\n\nand load back, my performance would drop significantly (&gt;0.02 on pubic LB), while the validation performance remains close, even if I run the identical model. Does anyone experiencing similar issue? Thank you!",
      "votes": 5
    },
    {
      "id": 320996,
      "postDate": "2018-04-30T10:59:22.373Z",
      "content": "<p>How do you save them?  If you write as csv then there may be rounding artifact.  Better save as binary (numpy, pickle, or feather files).</p>",
      "rawMarkdown": "How do you save them?  If you write as csv then there may be rounding artifact.  Better save as binary (numpy, pickle, or feather files).",
      "votes": 1,
      "replies": [
        {
          "id": 321417,
          "postDate": "2018-05-01T08:20:21.473Z",
          "content": "<p>Thanks @CPMP! I used .csv and then a very small set of 'click_id' has been rounded... Now i will use pickel, hope problem won't happen again.</p>",
          "rawMarkdown": "Thanks @CPMP! I used .csv and then a very small set of 'click_id' has been rounded... Now i will use pickel, hope problem won't happen again."
        },
        {
          "id": 323943,
          "postDate": "2018-05-06T18:39:21.303Z",
          "content": "<p>I personally switched to parquet during the competition. Quite fast to load and rather small in space.</p>",
          "rawMarkdown": "I personally switched to parquet during the competition. Quite fast to load and rather small in space."
        }
      ]
    },
    {
      "id": 323895,
      "postDate": "2018-05-06T15:45:22.653Z",
      "content": "<p>I wrote it several times, but the notebook you refer too saves features as csv.  This may truncate the floating point values.  It is much safer to use a binary format: numpy, pickle, or feather.</p>",
      "rawMarkdown": "I wrote it several times, but the notebook you refer too saves features as csv.  This may truncate the floating point values.  It is much safer to use a binary format: numpy, pickle, or feather.",
      "votes": 2
    },
    {
      "id": 322277,
      "postDate": "2018-05-02T16:57:52.627Z",
      "content": "<p>The same for me. I am also tried to debug those weeks, but this problem really confused me a lot actually, I thought my single lgb could achieve 9820+ at least, anyway, this is so weird for me. I am looking forward to us guys solution when this game end. : )</p>",
      "rawMarkdown": "The same for me. I am also tried to debug those weeks, but this problem really confused me a lot actually, I thought my single lgb could achieve 9820+ at least, anyway, this is so weird for me. I am looking forward to us guys solution when this game end. : )",
      "votes": 2
    },
    {
      "id": 320543,
      "postDate": "2018-04-29T02:18:04.463Z",
      "content": "<p>One watch-out is - make sure the order of features in the data frame you are scoring is identical to model design.</p>",
      "rawMarkdown": "One watch-out is - make sure the order of features in the data frame you are scoring is identical to model design.",
      "votes": 2,
      "replies": [
        {
          "id": 320553,
          "postDate": "2018-04-29T02:55:28.863Z",
          "content": "<p>Is it really necessary?? I have an altogether different order...</p>",
          "rawMarkdown": "Is it really necessary?? I have an altogether different order..."
        },
        {
          "id": 320558,
          "postDate": "2018-04-29T03:33:00.843Z",
          "content": "<p>Order of features do affect the performance of LGB</p>",
          "rawMarkdown": "Order of features do affect the performance of LGB",
          "votes": 6
        },
        {
          "id": 320564,
          "postDate": "2018-04-29T04:40:32.083Z",
          "content": "<p>Thanks for the confirmation.... </p>",
          "rawMarkdown": "Thanks for the confirmation.... "
        },
        {
          "id": 320822,
          "postDate": "2018-04-30T00:40:16.067Z",
          "content": "<p>@Pavel, @KALE, I just am re-running my validation scripts right now because of what I thought is this issue. I am so happy to see this confirmed here. With such a large dataset, some of us with not so powerful HW would have to do save and re-load a lot. Hence it is easy to forget such issues. Thanks </p>",
          "rawMarkdown": "@Pavel, @KALE, I just am re-running my validation scripts right now because of what I thought is this issue. I am so happy to see this confirmed here. With such a large dataset, some of us with not so powerful HW would have to do save and re-load a lot. Hence it is easy to forget such issues. Thanks "
        },
        {
          "id": 320892,
          "postDate": "2018-04-30T06:07:35.657Z",
          "content": "<p>@YaGana Sheriff-Hussaini  may I know how you rematch the index? </p>",
          "rawMarkdown": "@YaGana Sheriff-Hussaini  may I know how you rematch the index? "
        },
        {
          "id": 320963,
          "postDate": "2018-04-30T09:36:21.593Z",
          "content": "<p>@MengYe, in my case I have also added new features to my model, so saved it in the correct order this time.</p>\n\n<p>To answer your question, with pandas you just need to assign the columns in the order that you want. For example if df.columns = [\"b\", \"c\", \"d\", \"a\"], assign</p>\n\n<p>cols=[\"a\", \"b\", \"c\", \"d\"] then df.reindex(columns = cols).  Hope that helps.</p>",
          "rawMarkdown": "@MengYe, in my case I have also added new features to my model, so saved it in the correct order this time.\n\nTo answer your question, with pandas you just need to assign the columns in the order that you want. For example if df.columns = [\"b\", \"c\", \"d\", \"a\"], assign\n\ncols=[\"a\", \"b\", \"c\", \"d\"] then df.reindex(columns = cols).  Hope that helps."
        }
      ]
    },
    {
      "id": 322769,
      "postDate": "2018-05-03T16:02:17.483Z",
      "content": "<p>Cheng, have you found a solution to this problem? Save generated features in feather format helped? Thanks.</p>",
      "rawMarkdown": "Cheng, have you found a solution to this problem? Save generated features in feather format helped? Thanks.",
      "replies": [
        {
          "id": 323740,
          "postDate": "2018-05-06T03:33:47.713Z",
          "content": "<p>I always get significantly worse performance when I save the slicing column into local and save back. Give up on it now.</p>",
          "rawMarkdown": "I always get significantly worse performance when I save the slicing column into local and save back. Give up on it now."
        },
        {
          "id": 323875,
          "postDate": "2018-05-06T13:33:29.130Z",
          "content": "<p>I am also give up =_=. So, if just do FE and train and predict from scratch, it will be ok？</p>",
          "rawMarkdown": "I am also give up =_=. So, if just do FE and train and predict from scratch, it will be ok？"
        },
        {
          "id": 323892,
          "postDate": "2018-05-06T15:06:18.673Z",
          "content": "<p>Edited: moved to <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55581#323895\">another reply</a></p>",
          "rawMarkdown": "Edited: moved to [another reply][1]\n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55581#323895"
        },
        {
          "id": 323896,
          "postDate": "2018-05-06T15:47:28.130Z",
          "content": "<p>csv format.... tomorrow I will have my last try, thanks~</p>",
          "rawMarkdown": "csv format.... tomorrow I will have my last try, thanks~",
          "votes": 1
        },
        {
          "id": 323899,
          "postDate": "2018-05-06T15:55:41.837Z",
          "content": "<p>I hope for you it is the reason.  Maybe I'll regret it when you'll pass me on the LB... ;)</p>",
          "rawMarkdown": "I hope for you it is the reason.  Maybe I'll regret it when you'll pass me on the LB... ;)"
        },
        {
          "id": 323900,
          "postDate": "2018-05-06T16:01:36.557Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 320996,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-04-30T10:59:22.373000",
      "content": "<p>How do you save them?  If you write as csv then there may be rounding artifact.  Better save as binary (numpy, pickle, or feather files).</p>",
      "votes": 1,
      "replies": [
        {
          "id": 321417,
          "author_name": "Wenjie Bai",
          "author_url": "",
          "post_date": "2018-05-01T08:20:21.473000",
          "content": "<p>Thanks @CPMP! I used .csv and then a very small set of 'click_id' has been rounded... Now i will use pickel, hope problem won't happen again.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 323943,
          "author_name": "Dany Majard",
          "author_url": "",
          "post_date": "2018-05-06T18:39:21.303000",
          "content": "<p>I personally switched to parquet during the competition. Quite fast to load and rather small in space.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 323895,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-05-06T15:45:22.653000",
      "content": "<p>I wrote it several times, but the notebook you refer too saves features as csv.  This may truncate the floating point values.  It is much safer to use a binary format: numpy, pickle, or feather.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 322277,
      "author_name": "Fengari",
      "author_url": "",
      "post_date": "2018-05-02T16:57:52.627000",
      "content": "<p>The same for me. I am also tried to debug those weeks, but this problem really confused me a lot actually, I thought my single lgb could achieve 9820+ at least, anyway, this is so weird for me. I am looking forward to us guys solution when this game end. : )</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 320543,
      "author_name": "Pavel (Pasha) Gyrya",
      "author_url": "",
      "post_date": "2018-04-29T02:18:04.463000",
      "content": "<p>One watch-out is - make sure the order of features in the data frame you are scoring is identical to model design.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 320553,
          "author_name": "Samrat Pandiri",
          "author_url": "",
          "post_date": "2018-04-29T02:55:28.863000",
          "content": "<p>Is it really necessary?? I have an altogether different order...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320558,
          "author_name": "KALE",
          "author_url": "",
          "post_date": "2018-04-29T03:33:00.843000",
          "content": "<p>Order of features do affect the performance of LGB</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 320564,
          "author_name": "Samrat Pandiri",
          "author_url": "",
          "post_date": "2018-04-29T04:40:32.083000",
          "content": "<p>Thanks for the confirmation.... </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320822,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-04-30T00:40:16.067000",
          "content": "<p>@Pavel, @KALE, I just am re-running my validation scripts right now because of what I thought is this issue. I am so happy to see this confirmed here. With such a large dataset, some of us with not so powerful HW would have to do save and re-load a lot. Hence it is easy to forget such issues. Thanks </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320892,
          "author_name": "MengYe",
          "author_url": "",
          "post_date": "2018-04-30T06:07:35.657000",
          "content": "<p>@YaGana Sheriff-Hussaini  may I know how you rematch the index? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320963,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-04-30T09:36:21.593000",
          "content": "<p>@MengYe, in my case I have also added new features to my model, so saved it in the correct order this time.</p>\n\n<p>To answer your question, with pandas you just need to assign the columns in the order that you want. For example if df.columns = [\"b\", \"c\", \"d\", \"a\"], assign</p>\n\n<p>cols=[\"a\", \"b\", \"c\", \"d\"] then df.reindex(columns = cols).  Hope that helps.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 322769,
      "author_name": "Paulo Felipe",
      "author_url": "",
      "post_date": "2018-05-03T16:02:17.483000",
      "content": "<p>Cheng, have you found a solution to this problem? Save generated features in feather format helped? Thanks.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 323740,
          "author_name": "Cheng",
          "author_url": "",
          "post_date": "2018-05-06T03:33:47.713000",
          "content": "<p>I always get significantly worse performance when I save the slicing column into local and save back. Give up on it now.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 323875,
          "author_name": "Fengari",
          "author_url": "",
          "post_date": "2018-05-06T13:33:29.130000",
          "content": "<p>I am also give up =_=. So, if just do FE and train and predict from scratch, it will be ok？</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 323892,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-06T15:06:18.673000",
          "content": "<p>Edited: moved to <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55581#323895\">another reply</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 323896,
          "author_name": "Fengari",
          "author_url": "",
          "post_date": "2018-05-06T15:47:28.130000",
          "content": "<p>csv format.... tomorrow I will have my last try, thanks~</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 323899,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-06T15:55:41.837000",
          "content": "<p>I hope for you it is the reason.  Maybe I'll regret it when you'll pass me on the LB... ;)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 323900,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-06T16:01:36.557000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "320510": "I am experiencing a very strange bug: when I save the generated feature to local disk as in\n\nhttps://www.kaggle.com/bk0000/non-blending-lightgbm-model-lb-0-977\n\nand load back, my performance would drop significantly (&gt;0.02 on pubic LB), while the validation performance remains close, even if I run the identical model. Does anyone experiencing similar issue? Thank you!",
    "320996": "How do you save them?  If you write as csv then there may be rounding artifact.  Better save as binary (numpy, pickle, or feather files).",
    "323895": "I wrote it several times, but the notebook you refer too saves features as csv.  This may truncate the floating point values.  It is much safer to use a binary format: numpy, pickle, or feather.",
    "322277": "The same for me. I am also tried to debug those weeks, but this problem really confused me a lot actually, I thought my single lgb could achieve 9820+ at least, anyway, this is so weird for me. I am looking forward to us guys solution when this game end. : )",
    "320543": "One watch-out is - make sure the order of features in the data frame you are scoring is identical to model design.",
    "322769": "Cheng, have you found a solution to this problem? Save generated features in feather format helped? Thanks."
  }
}