{
  "id": 386473,
  "title": "Maybe Feature Engineering is more important than training and tuning model",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/386473",
  "author_name": "",
  "post_date": "2023-02-13T09:51:19.380345100Z",
  "votes": 7,
  "comment_count": 8,
  "views": 0,
  "content": "<h3><a href=\"https://www.kaggle.com/mohammad2012191\" target=\"_blank\">@mohammad2012191</a> 's baseline lightGBM (LB 0.670)</h3>\n<p><a href=\"https://www.kaggle.com/code/mohammad2012191/lgbm-early-stopping-lb-0-670\" target=\"_blank\">https://www.kaggle.com/code/mohammad2012191/lgbm-early-stopping-lb-0-670</a><br>\nFE:</p>\n<pre><code>CATS = [, ,, , ]\nNUMS = [,,,, , \n        , , ]\n</code></pre>\n<h3><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>  got LB 0.676 ( +0.006 improvement) with more features</h3>\n<p><a href=\"https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-676\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-676</a></p>\n<pre><code>CATS = [, , , ]\nNUMS = [,,,, , \n        , , ]\n\n\nEVENTS = [,,,,\n          ,,,,\n          ]\n</code></pre>\n<p><strong>These new features are results of one-hot of event_name</strong></p>\n<h3><a href=\"https://www.kaggle.com/starfalllover\" target=\"_blank\">@starfalllover</a> tuned the parameters of model by standing on the shoudlers of <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> (LB 0.678)</h3>\n<p><a href=\"https://www.kaggle.com/code/starfalllover/xgboost-baseline-0-678\" target=\"_blank\">https://www.kaggle.com/code/starfalllover/xgboost-baseline-0-678</a><br>\ngot a +0.002 improvement</p>\n<p>so, maybe pay more attentions to find out new features is more efficiently than tuning model.</p>",
  "messages": [
    {
      "id": "2142056",
      "postDate": "02/13/2023 09:51:19",
      "content": "<h3><a href=\"https://www.kaggle.com/mohammad2012191\" target=\"_blank\">@mohammad2012191</a> 's baseline lightGBM (LB 0.670)</h3>\n<p><a href=\"https://www.kaggle.com/code/mohammad2012191/lgbm-early-stopping-lb-0-670\" target=\"_blank\">https://www.kaggle.com/code/mohammad2012191/lgbm-early-stopping-lb-0-670</a><br>\nFE:</p>\n<pre><code>CATS = [, ,, , ]\nNUMS = [,,,, , \n        , , ]\n</code></pre>\n<h3><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>  got LB 0.676 ( +0.006 improvement) with more features</h3>\n<p><a href=\"https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-676\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-676</a></p>\n<pre><code>CATS = [, , , ]\nNUMS = [,,,, , \n        , , ]\n\n\nEVENTS = [,,,,\n          ,,,,\n          ]\n</code></pre>\n<p><strong>These new features are results of one-hot of event_name</strong></p>\n<h3><a href=\"https://www.kaggle.com/starfalllover\" target=\"_blank\">@starfalllover</a> tuned the parameters of model by standing on the shoudlers of <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> (LB 0.678)</h3>\n<p><a href=\"https://www.kaggle.com/code/starfalllover/xgboost-baseline-0-678\" target=\"_blank\">https://www.kaggle.com/code/starfalllover/xgboost-baseline-0-678</a><br>\ngot a +0.002 improvement</p>\n<p>so, maybe pay more attentions to find out new features is more efficiently than tuning model.</p>",
      "rawMarkdown": "### @mohammad2012191 's baseline lightGBM (LB 0.670)\nhttps://www.kaggle.com/code/mohammad2012191/lgbm-early-stopping-lb-0-670\nFE:\n```python\nCATS = ['event_name', 'name','fqid', 'room_fqid', 'text_fqid']\nNUMS = ['elapsed_time','level','page','room_coor_x', 'room_coor_y', \n        'screen_coor_x', 'screen_coor_y', 'hover_duration']\n```\n\n\n### @cdeotte  got LB 0.676 ( +0.006 improvement) with more features\nhttps://www.kaggle.com/code/cdeotte/xgboost-baseline-0-676\n```python\nCATS = ['event_name', 'fqid', 'room_fqid', 'text']\nNUMS = ['elapsed_time','level','page','room_coor_x', 'room_coor_y', \n        'screen_coor_x', 'screen_coor_y', 'hover_duration']\n\n# https://www.kaggle.com/code/kimtaehun/lightgbm-baseline-with-aggregated-log-data\nEVENTS = ['navigate_click','person_click','cutscene_click','object_click',\n          'map_hover','notification_click','map_click','observation_click',\n          'checkpoint']\n```\n**These new features are results of one-hot of event_name**\n\n### @starfalllover tuned the parameters of model by standing on the shoudlers of @cdeotte (LB 0.678)\nhttps://www.kaggle.com/code/starfalllover/xgboost-baseline-0-678\ngot a +0.002 improvement\n\nso, maybe pay more attentions to find out new features is more efficiently than tuning model.",
      "votes": null
    },
    {
      "id": "2142117",
      "postDate": "02/13/2023 10:27:16",
      "content": "<p>Yeap, not just for this, but performing good EDA to properly do feature engineering is more important IMO than training and tuning model</p>",
      "rawMarkdown": "Yeap, not just for this, but performing good EDA to properly do feature engineering is more important IMO than training and tuning model",
      "votes": null
    },
    {
      "id": "2142174",
      "postDate": "02/13/2023 11:07:39",
      "content": "<p>yup it seems the way to go for xgboost…. at least until LSTM come into play. </p>",
      "rawMarkdown": "yup it seems the way to go for xgboost.... at least until LSTM come into play.",
      "votes": null
    },
    {
      "id": "2143034",
      "postDate": "02/14/2023 01:40:29",
      "content": "<p>Yes, you are right.<br>\nI have done anther ablation experiment based on <a href=\"https://www.kaggle.com/mohammad2012191\" target=\"_blank\">@mohammad2012191</a>'s code, but only use the following features and LB &amp; CV both are 0.673. </p>\n<pre><code>EVENTS = [,,,,\n          ,,,,\n          ]\n ():\n\n    dfs = []\n     c  EVENTS: \n        train[c] = (train.event_name == c).astype()\n     c  EVENTS + []:\n        tmp = train.groupby([,])[c].agg()\n        tmp.name = tmp.name + \n        dfs.append(tmp)\n    train = train.drop(EVENTS,axis=)\n\n    df = pd.concat(dfs,axis=)\n    df = df.fillna(-)\n    df = df.reset_index()\n    df = df.set_index()\n     df\n</code></pre>\n<p>so, I suppose event_name is the main point for this competition.</p>",
      "rawMarkdown": "Yes, you are right.\nI have done anther ablation experiment based on @mohammad2012191's code, but only use the following features and LB & CV both are 0.673. \n```python\nEVENTS = ['navigate_click','person_click','cutscene_click','object_click',\n          'map_hover','notification_click','map_click','observation_click',\n          'checkpoint']\ndef feature_engineer(train):\n    \n    dfs = []\n    for c in EVENTS: \n        train[c] = (train.event_name == c).astype('int8')\n    for c in EVENTS + ['elapsed_time']:\n        tmp = train.groupby(['session_id','level_group'])[c].agg('sum')\n        tmp.name = tmp.name + '_sum'\n        dfs.append(tmp)\n    train = train.drop(EVENTS,axis=1)\n        \n    df = pd.concat(dfs,axis=1)\n    df = df.fillna(-1)\n    df = df.reset_index()\n    df = df.set_index('session_id')\n    return df\n```\nso, I suppose event_name is the main point for this competition.",
      "votes": null
    },
    {
      "id": "2143078",
      "postDate": "02/14/2023 02:36:15",
      "content": "<p>yeah, maybe we can think about how to present current data now, so that we can import more complex model, like LSTM or graph neural network, into this game.</p>",
      "rawMarkdown": "yeah, maybe we can think about how to present current data now, so that we can import more complex model, like LSTM or graph neural network, into this game.",
      "votes": null
    },
    {
      "id": "2143991",
      "postDate": "02/14/2023 17:02:34",
      "content": "<p>Hmm interesting analysis. Thanks for showing the code here too!</p>",
      "rawMarkdown": "Hmm interesting analysis. Thanks for showing the code here too!",
      "votes": null
    },
    {
      "id": "2145281",
      "postDate": "02/14/2023 22:31:49",
      "content": "<p>Interesting.Thanks.</p>",
      "rawMarkdown": "Interesting.Thanks.",
      "votes": null
    },
    {
      "id": "2147160",
      "postDate": "02/16/2023 12:52:20",
      "content": "<p>I think time is better spent on engenering features than tuning hyperparams.</p>",
      "rawMarkdown": "I think time is better spent on engenering features than tuning hyperparams.",
      "votes": null
    },
    {
      "id": "2147403",
      "postDate": "02/16/2023 15:37:34",
      "content": "<p><a href=\"https://www.kaggle.com/mengvision\" target=\"_blank\">@mengvision</a> Thanks for the proof. Yes, the quality of the data is more important than the model itself. The model can only as good as the input (engineered) data.</p>",
      "rawMarkdown": "mengvision Thanks for the proof. Yes, the quality of the data is more important than the model itself. The model can only as good as the input (engineered) data.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2142117,
      "author_name": "fahaddalwai",
      "author_url": "",
      "post_date": "02/13/2023 10:27:16",
      "content": "<p>Yeap, not just for this, but performing good EDA to properly do feature engineering is more important IMO than training and tuning model</p>",
      "votes": null,
      "replies": [
        {
          "id": 2143034,
          "author_name": "mengvision",
          "author_url": "",
          "post_date": "02/14/2023 01:40:29",
          "content": "<p>Yes, you are right.<br>\nI have done anther ablation experiment based on <a href=\"https://www.kaggle.com/mohammad2012191\" target=\"_blank\">@mohammad2012191</a>'s code, but only use the following features and LB &amp; CV both are 0.673. </p>\n<pre><code>EVENTS = [,,,,\n          ,,,,\n          ]\n ():\n\n    dfs = []\n     c  EVENTS: \n        train[c] = (train.event_name == c).astype()\n     c  EVENTS + []:\n        tmp = train.groupby([,])[c].agg()\n        tmp.name = tmp.name + \n        dfs.append(tmp)\n    train = train.drop(EVENTS,axis=)\n\n    df = pd.concat(dfs,axis=)\n    df = df.fillna(-)\n    df = df.reset_index()\n    df = df.set_index()\n     df\n</code></pre>\n<p>so, I suppose event_name is the main point for this competition.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2143991,
              "author_name": "fahaddalwai",
              "author_url": "",
              "post_date": "02/14/2023 17:02:34",
              "content": "<p>Hmm interesting analysis. Thanks for showing the code here too!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2142174,
      "author_name": "lucasmorin",
      "author_url": "",
      "post_date": "02/13/2023 11:07:39",
      "content": "<p>yup it seems the way to go for xgboost…. at least until LSTM come into play. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2143078,
          "author_name": "mengvision",
          "author_url": "",
          "post_date": "02/14/2023 02:36:15",
          "content": "<p>yeah, maybe we can think about how to present current data now, so that we can import more complex model, like LSTM or graph neural network, into this game.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2145281,
      "author_name": "tomonorisasaki",
      "author_url": "",
      "post_date": "02/14/2023 22:31:49",
      "content": "<p>Interesting.Thanks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2147160,
      "author_name": "simonveitner",
      "author_url": "",
      "post_date": "02/16/2023 12:52:20",
      "content": "<p>I think time is better spent on engenering features than tuning hyperparams.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2147403,
      "author_name": "minhbtnguyen",
      "author_url": "",
      "post_date": "02/16/2023 15:37:34",
      "content": "<p><a href=\"https://www.kaggle.com/mengvision\" target=\"_blank\">@mengvision</a> Thanks for the proof. Yes, the quality of the data is more important than the model itself. The model can only as good as the input (engineered) data.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2142056": "### @mohammad2012191 's baseline lightGBM (LB 0.670)\nhttps://www.kaggle.com/code/mohammad2012191/lgbm-early-stopping-lb-0-670\nFE:\n```python\nCATS = ['event_name', 'name','fqid', 'room_fqid', 'text_fqid']\nNUMS = ['elapsed_time','level','page','room_coor_x', 'room_coor_y', \n        'screen_coor_x', 'screen_coor_y', 'hover_duration']\n```\n\n\n### @cdeotte  got LB 0.676 ( +0.006 improvement) with more features\nhttps://www.kaggle.com/code/cdeotte/xgboost-baseline-0-676\n```python\nCATS = ['event_name', 'fqid', 'room_fqid', 'text']\nNUMS = ['elapsed_time','level','page','room_coor_x', 'room_coor_y', \n        'screen_coor_x', 'screen_coor_y', 'hover_duration']\n\n# https://www.kaggle.com/code/kimtaehun/lightgbm-baseline-with-aggregated-log-data\nEVENTS = ['navigate_click','person_click','cutscene_click','object_click',\n          'map_hover','notification_click','map_click','observation_click',\n          'checkpoint']\n```\n**These new features are results of one-hot of event_name**\n\n### @starfalllover tuned the parameters of model by standing on the shoudlers of @cdeotte (LB 0.678)\nhttps://www.kaggle.com/code/starfalllover/xgboost-baseline-0-678\ngot a +0.002 improvement\n\nso, maybe pay more attentions to find out new features is more efficiently than tuning model.",
    "2142117": "Yeap, not just for this, but performing good EDA to properly do feature engineering is more important IMO than training and tuning model",
    "2142174": "yup it seems the way to go for xgboost.... at least until LSTM come into play.",
    "2143034": "Yes, you are right.\nI have done anther ablation experiment based on @mohammad2012191's code, but only use the following features and LB & CV both are 0.673. \n```python\nEVENTS = ['navigate_click','person_click','cutscene_click','object_click',\n          'map_hover','notification_click','map_click','observation_click',\n          'checkpoint']\ndef feature_engineer(train):\n    \n    dfs = []\n    for c in EVENTS: \n        train[c] = (train.event_name == c).astype('int8')\n    for c in EVENTS + ['elapsed_time']:\n        tmp = train.groupby(['session_id','level_group'])[c].agg('sum')\n        tmp.name = tmp.name + '_sum'\n        dfs.append(tmp)\n    train = train.drop(EVENTS,axis=1)\n        \n    df = pd.concat(dfs,axis=1)\n    df = df.fillna(-1)\n    df = df.reset_index()\n    df = df.set_index('session_id')\n    return df\n```\nso, I suppose event_name is the main point for this competition.",
    "2143078": "yeah, maybe we can think about how to present current data now, so that we can import more complex model, like LSTM or graph neural network, into this game.",
    "2143991": "Hmm interesting analysis. Thanks for showing the code here too!",
    "2145281": "Interesting.Thanks.",
    "2147160": "I think time is better spent on engenering features than tuning hyperparams.",
    "2147403": "mengvision Thanks for the proof. Yes, the quality of the data is more important than the model itself. The model can only as good as the input (engineered) data."
  },
  "source": "meta"
}