{
  "id": 195125,
  "title": "merge feature to test_df cost too many time",
  "url": "/competitions/riiid-test-answer-prediction/discussion/195125",
  "author_name": "huangtaogan",
  "post_date": "2020-11-03T16:29:51.393000",
  "votes": 0,
  "comment_count": 4,
  "views": 0,
  "content": "<p>the notebook should finish the prediction in 9h, and I found it waste too many time to merge feature to test_df. In my notebook, the performance for provided 120 sample is：<br>\n4.091254999999997s：whole time the cell which make prediction cost<br>\n3.030566999999998s：whole time to merge feature <br>\n0.017841000000004215s：whole time LGBM with 1000rounds make prediction</p>\n<p>any idea to improve the rate of merge feature? by the way, I use pandas.merge to do it, I try to set_index and use join, but it works nothing</p>",
  "messages": [
    {
      "id": 1068730,
      "postDate": "2020-11-03T17:22:54.243Z",
      "content": "<p>numpy is faster than pandas</p>",
      "rawMarkdown": "numpy is faster than pandas",
      "votes": -1
    },
    {
      "id": 1068679,
      "postDate": "2020-11-03T16:29:51.393Z",
      "content": "<p>the notebook should finish the prediction in 9h, and I found it waste too many time to merge feature to test_df. In my notebook, the performance for provided 120 sample is：<br>\n4.091254999999997s：whole time the cell which make prediction cost<br>\n3.030566999999998s：whole time to merge feature <br>\n0.017841000000004215s：whole time LGBM with 1000rounds make prediction</p>\n<p>any idea to improve the rate of merge feature? by the way, I use pandas.merge to do it, I try to set_index and use join, but it works nothing</p>",
      "rawMarkdown": "the notebook should finish the prediction in 9h, and I found it waste too many time to merge feature to test_df. In my notebook, the performance for provided 120 sample is：\n4.091254999999997s：whole time the cell which make prediction cost\n3.030566999999998s：whole time to merge feature \n0.017841000000004215s：whole time LGBM with 1000rounds make prediction\n\nany idea to improve the rate of merge feature? by the way, I use pandas.merge to do it, I try to set_index and use join, but it works nothing"
    },
    {
      "id": 1076265,
      "postDate": "2020-11-12T11:37:19.327Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1070006,
      "postDate": "2020-11-05T08:16:18.403Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1070359,
          "postDate": "2020-11-05T16:58:17.897Z",
          "content": "<p>☹️no idea yet</p>",
          "rawMarkdown": "☹️no idea yet"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1068730,
      "author_name": "Pavel Orlov",
      "author_url": "",
      "post_date": "2020-11-03T17:22:54.243000",
      "content": "<p>numpy is faster than pandas</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 1076265,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-11-12T11:37:19.327000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1070006,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-11-05T08:16:18.403000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1070359,
          "author_name": "huangtaogan",
          "author_url": "",
          "post_date": "2020-11-05T16:58:17.897000",
          "content": "<p>☹️no idea yet</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1068730": "numpy is faster than pandas",
    "1068679": "the notebook should finish the prediction in 9h, and I found it waste too many time to merge feature to test_df. In my notebook, the performance for provided 120 sample is：\n4.091254999999997s：whole time the cell which make prediction cost\n3.030566999999998s：whole time to merge feature \n0.017841000000004215s：whole time LGBM with 1000rounds make prediction\n\nany idea to improve the rate of merge feature? by the way, I use pandas.merge to do it, I try to set_index and use join, but it works nothing",
    "1076265": "",
    "1070006": ""
  }
}