{
  "id": 209583,
  "title": "LB 0.796 single LightGBM model with Train/Infer code[updating some ideas]",
  "url": "/competitions/riiid-test-answer-prediction/discussion/209583",
  "author_name": "",
  "post_date": "2021-01-08T00:19:19.196208500Z",
  "votes": 22,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Firstly, congratulations to every winner and respect to everyone who works hard, just sharing my LB 0.796 single LightGBM model kernel  <a href=\"https://www.kaggle.com/a763337092/lgb1215?scriptVersionId=49450563\" target=\"_blank\">https://www.kaggle.com/a763337092/lgb1215?scriptVersionId=49450563</a></p>\n<p>For there is almost no high score training code sharing, I'd like to share this code with you.</p>\n<p>In feature engineering, the most important discoveries are:</p>\n<ol>\n<li>user_content_mean_mean: the avg of every user's content_id answer_correctly mean. My local cv improves 0.07 by adding this feature.</li>\n<li>user_last_timespan: the timespan from last to this action of every user, it improves nearly 0.01.</li>\n<li>user_content_cnt: looping count to get every user’s groupby(['user_id', 'content_id'])['content_id'].count().</li>\n</ol>\n<p>As saying, it contains both training and infering part in this code. When 'OFFLINE  = True', it is for training and for infering when 'OFFLINE  = False'. You can run training locally and run infering in kernel.</p>\n<p>Good Luck!</p>",
  "messages": [
    {
      "id": "1143544",
      "postDate": "01/08/2021 00:19:19",
      "content": "<p>Firstly, congratulations to every winner and respect to everyone who works hard, just sharing my LB 0.796 single LightGBM model kernel  <a href=\"https://www.kaggle.com/a763337092/lgb1215?scriptVersionId=49450563\" target=\"_blank\">https://www.kaggle.com/a763337092/lgb1215?scriptVersionId=49450563</a></p>\n<p>For there is almost no high score training code sharing, I'd like to share this code with you.</p>\n<p>In feature engineering, the most important discoveries are:</p>\n<ol>\n<li>user_content_mean_mean: the avg of every user's content_id answer_correctly mean. My local cv improves 0.07 by adding this feature.</li>\n<li>user_last_timespan: the timespan from last to this action of every user, it improves nearly 0.01.</li>\n<li>user_content_cnt: looping count to get every user’s groupby(['user_id', 'content_id'])['content_id'].count().</li>\n</ol>\n<p>As saying, it contains both training and infering part in this code. When 'OFFLINE  = True', it is for training and for infering when 'OFFLINE  = False'. You can run training locally and run infering in kernel.</p>\n<p>Good Luck!</p>",
      "rawMarkdown": "Firstly, congratulations to every winner and respect to everyone who works hard, just sharing my LB 0.796 single LightGBM model kernel  https://www.kaggle.com/a763337092/lgb1215?scriptVersionId=49450563\n\nFor there is almost no high score training code sharing, I'd like to share this code with you.\n\nIn feature engineering, the most important discoveries are:\n1. user_content_mean_mean: the avg of every user's content_id answer_correctly mean. My local cv improves 0.07 by adding this feature.\n2. user_last_timespan: the timespan from last to this action of every user, it improves nearly 0.01.\n3. user_content_cnt: looping count to get every user’s groupby(['user_id', 'content_id'])['content_id'].count().\n\nAs saying, it contains both training and infering part in this code. When 'OFFLINE  = True', it is for training and for infering when 'OFFLINE  = False'. You can run training locally and run infering in kernel.\n\nGood Luck!",
      "votes": null
    },
    {
      "id": "1143547",
      "postDate": "01/08/2021 00:20:33",
      "content": "<p><a href=\"https://www.kaggle.com/a763337092\" target=\"_blank\">@a763337092</a> great work. Would you be able to add a feature importance plot? I saw that you logged the increases in score from the features you added but would still be cool to see that plot :)</p>",
      "rawMarkdown": "a763337092 great work. Would you be able to add a feature importance plot? I saw that you logged the increases in score from the features you added but would still be cool to see that plot :)",
      "votes": null
    },
    {
      "id": "1143554",
      "postDate": "01/08/2021 00:24:02",
      "content": "<p>OK, I‘ll rerun a 1/10 sampling and add it</p>",
      "rawMarkdown": "OK, I‘ll rerun a 1/10 sampling and add it",
      "votes": null
    },
    {
      "id": "1143580",
      "postDate": "01/08/2021 00:42:07",
      "content": "<p>Thank you for sharing your great LGBM approach! </p>",
      "rawMarkdown": "Thank you for sharing your great LGBM approach!",
      "votes": null
    },
    {
      "id": "1143648",
      "postDate": "01/08/2021 01:39:40",
      "content": "<p>👍👍👍👍👍👍👍👍👍👍👍👍👍👍</p>",
      "rawMarkdown": "👍👍👍👍👍👍👍👍👍👍👍👍👍👍",
      "votes": null
    },
    {
      "id": "1143701",
      "postDate": "01/08/2021 02:31:39",
      "content": "<p>Carrrry me Bao👀👀</p>",
      "rawMarkdown": "Carrrry me Bao👀👀",
      "votes": null
    },
    {
      "id": "1143702",
      "postDate": "01/08/2021 02:33:02",
      "content": "<p>My pleasure😃</p>",
      "rawMarkdown": "My pleasure😃",
      "votes": null
    },
    {
      "id": "1144727",
      "postDate": "01/08/2021 16:25:00",
      "content": "<p>nice! a great job!</p>",
      "rawMarkdown": "nice! a great job!",
      "votes": null
    },
    {
      "id": "1147616",
      "postDate": "01/10/2021 15:58:16",
      "content": "<p>Thanks for sharing! sorry to bother, if train in kernel how many data did you use to avoid memory error?</p>",
      "rawMarkdown": "Thanks for sharing! sorry to bother, if train in kernel how many data did you use to avoid memory error?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1143547,
      "author_name": "rdizzl3",
      "author_url": "",
      "post_date": "01/08/2021 00:20:33",
      "content": "<p><a href=\"https://www.kaggle.com/a763337092\" target=\"_blank\">@a763337092</a> great work. Would you be able to add a feature importance plot? I saw that you logged the increases in score from the features you added but would still be cool to see that plot :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1143554,
          "author_name": "a763337092",
          "author_url": "",
          "post_date": "01/08/2021 00:24:02",
          "content": "<p>OK, I‘ll rerun a 1/10 sampling and add it</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1143580,
      "author_name": "sishihara",
      "author_url": "",
      "post_date": "01/08/2021 00:42:07",
      "content": "<p>Thank you for sharing your great LGBM approach! </p>",
      "votes": null,
      "replies": [
        {
          "id": 1143702,
          "author_name": "a763337092",
          "author_url": "",
          "post_date": "01/08/2021 02:33:02",
          "content": "<p>My pleasure😃</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1143648,
      "author_name": "baomengjiao",
      "author_url": "",
      "post_date": "01/08/2021 01:39:40",
      "content": "<p>👍👍👍👍👍👍👍👍👍👍👍👍👍👍</p>",
      "votes": null,
      "replies": [
        {
          "id": 1143701,
          "author_name": "a763337092",
          "author_url": "",
          "post_date": "01/08/2021 02:31:39",
          "content": "<p>Carrrry me Bao👀👀</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1144727,
      "author_name": "zjjszj2",
      "author_url": "",
      "post_date": "01/08/2021 16:25:00",
      "content": "<p>nice! a great job!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1147616,
      "author_name": "zjjszj2",
      "author_url": "",
      "post_date": "01/10/2021 15:58:16",
      "content": "<p>Thanks for sharing! sorry to bother, if train in kernel how many data did you use to avoid memory error?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1143544": "Firstly, congratulations to every winner and respect to everyone who works hard, just sharing my LB 0.796 single LightGBM model kernel  https://www.kaggle.com/a763337092/lgb1215?scriptVersionId=49450563\n\nFor there is almost no high score training code sharing, I'd like to share this code with you.\n\nIn feature engineering, the most important discoveries are:\n1. user_content_mean_mean: the avg of every user's content_id answer_correctly mean. My local cv improves 0.07 by adding this feature.\n2. user_last_timespan: the timespan from last to this action of every user, it improves nearly 0.01.\n3. user_content_cnt: looping count to get every user’s groupby(['user_id', 'content_id'])['content_id'].count().\n\nAs saying, it contains both training and infering part in this code. When 'OFFLINE  = True', it is for training and for infering when 'OFFLINE  = False'. You can run training locally and run infering in kernel.\n\nGood Luck!",
    "1143547": "a763337092 great work. Would you be able to add a feature importance plot? I saw that you logged the increases in score from the features you added but would still be cool to see that plot :)",
    "1143554": "OK, I‘ll rerun a 1/10 sampling and add it",
    "1143580": "Thank you for sharing your great LGBM approach!",
    "1143648": "👍👍👍👍👍👍👍👍👍👍👍👍👍👍",
    "1143701": "Carrrry me Bao👀👀",
    "1143702": "My pleasure😃",
    "1144727": "nice! a great job!",
    "1147616": "Thanks for sharing! sorry to bother, if train in kernel how many data did you use to avoid memory error?"
  },
  "source": "meta"
}