{
  "id": 246675,
  "title": "Any tutorials for feature engineering for the new",
  "url": "/competitions/mlb-player-digital-engagement-forecasting/discussion/246675",
  "author_name": "",
  "post_date": "2021-06-16T11:38:28.114111400Z",
  "votes": 6,
  "comment_count": 6,
  "views": 0,
  "content": "<p>It will be grateful for providing some efficient pipelines to start a feature engineering competition. 🤓</p>",
  "messages": [
    {
      "id": "1351539",
      "postDate": "06/16/2021 11:38:28",
      "content": "<p>It will be grateful for providing some efficient pipelines to start a feature engineering competition. 🤓</p>",
      "rawMarkdown": "It will be grateful for providing some efficient pipelines to start a feature engineering competition. 🤓",
      "votes": null
    },
    {
      "id": "1354566",
      "postDate": "06/17/2021 17:01:30",
      "content": "<p>There are all kinds of levels to feature engineering from the most basic to really advanced. I'm not sure what you are looking for exactly. </p>",
      "rawMarkdown": "There are all kinds of levels to feature engineering from the most basic to really advanced. I'm not sure what you are looking for exactly.",
      "votes": null
    },
    {
      "id": "1356300",
      "postDate": "06/18/2021 22:03:05",
      "content": "<p>Most of feature engineering is basically common sense and creativity. Come up with ideas first, then figure out how to code them. Many ideas can be done with simple pandas functionality like group by and merge. Others might be trickier coding problems. I'd start with basic 1-1 transformations like date -&gt; year, month, day-of-month and day-of-week. Then I'd do basic summary stuff on all the corresponding tables, and merge those to your training set. For example, a flag for \"did they play yesterday\" (left join games data on playerid, and then use .isna()). Another example is Players' teams' number of followers (join roster to get team id, then join teamFollowers to that). Finally, I'd try some simple time-dependent features, like basically any other feature or target from a day earlier. One way to do this is to copy your features/targets table, replace the date with (date + 1), rename columns \"…_lagged\" and merge that back onto your training data set.</p>\n<p>Good luck!</p>",
      "rawMarkdown": "Most of feature engineering is basically common sense and creativity. Come up with ideas first, then figure out how to code them. Many ideas can be done with simple pandas functionality like group by and merge. Others might be trickier coding problems. I'd start with basic 1-1 transformations like date -> year, month, day-of-month and day-of-week. Then I'd do basic summary stuff on all the corresponding tables, and merge those to your training set. For example, a flag for \"did they play yesterday\" (left join games data on playerid, and then use .isna()). Another example is Players' teams' number of followers (join roster to get team id, then join teamFollowers to that). Finally, I'd try some simple time-dependent features, like basically any other feature or target from a day earlier. One way to do this is to copy your features/targets table, replace the date with (date + 1), rename columns \"..._lagged\" and merge that back onto your training data set.\n\nGood luck!",
      "votes": null
    },
    {
      "id": "1356382",
      "postDate": "06/19/2021 01:42:15",
      "content": "<p>I recommend looking at these two notebooks and building on similar features. Good luck!  Make sure you up-vote any code you reference/use so they get credit.</p>\n<p><a href=\"https://www.kaggle.com/ulrich07/mlb-ann-with-lags-tf-keras\" target=\"_blank\">https://www.kaggle.com/ulrich07/mlb-ann-with-lags-tf-keras</a></p>\n<p><a href=\"https://www.kaggle.com/columbia2131/mlb-lightgbm-starter-dataset-code-en-ja\" target=\"_blank\">https://www.kaggle.com/columbia2131/mlb-lightgbm-starter-dataset-code-en-ja</a></p>\n<p><a href=\"https://www.kaggle.com/ulrich07\" target=\"_blank\">@ulrich07</a> <a href=\"https://www.kaggle.com/columbia2131\" target=\"_blank\">@columbia2131</a></p>",
      "rawMarkdown": "I recommend looking at these two notebooks and building on similar features. Good luck!  Make sure you up-vote any code you reference/use so they get credit.\n\nhttps://www.kaggle.com/ulrich07/mlb-ann-with-lags-tf-keras\n\nhttps://www.kaggle.com/columbia2131/mlb-lightgbm-starter-dataset-code-en-ja\n\n@ulrich07 @columbia2131",
      "votes": null
    },
    {
      "id": "1356390",
      "postDate": "06/19/2021 02:00:02",
      "content": "<p>thx for your reply!</p>",
      "rawMarkdown": "thx for your reply!",
      "votes": null
    },
    {
      "id": "1356392",
      "postDate": "06/19/2021 02:00:43",
      "content": "<p>thx, this will be a great help!</p>",
      "rawMarkdown": "thx, this will be a great help!",
      "votes": null
    },
    {
      "id": "1392472",
      "postDate": "07/18/2021 18:13:37",
      "content": "<p>Give a try with <code>tsfresh</code> package if you need time-series features <a href=\"https://tsfresh.readthedocs.io/en/latest/\" target=\"_blank\">https://tsfresh.readthedocs.io/en/latest/</a></p>",
      "rawMarkdown": "Give a try with `tsfresh` package if you need time-series features https://tsfresh.readthedocs.io/en/latest/",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1354566,
      "author_name": "crained",
      "author_url": "",
      "post_date": "06/17/2021 17:01:30",
      "content": "<p>There are all kinds of levels to feature engineering from the most basic to really advanced. I'm not sure what you are looking for exactly. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1356300,
      "author_name": "paulfornia",
      "author_url": "",
      "post_date": "06/18/2021 22:03:05",
      "content": "<p>Most of feature engineering is basically common sense and creativity. Come up with ideas first, then figure out how to code them. Many ideas can be done with simple pandas functionality like group by and merge. Others might be trickier coding problems. I'd start with basic 1-1 transformations like date -&gt; year, month, day-of-month and day-of-week. Then I'd do basic summary stuff on all the corresponding tables, and merge those to your training set. For example, a flag for \"did they play yesterday\" (left join games data on playerid, and then use .isna()). Another example is Players' teams' number of followers (join roster to get team id, then join teamFollowers to that). Finally, I'd try some simple time-dependent features, like basically any other feature or target from a day earlier. One way to do this is to copy your features/targets table, replace the date with (date + 1), rename columns \"…_lagged\" and merge that back onto your training data set.</p>\n<p>Good luck!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1356390,
          "author_name": "vanle73",
          "author_url": "",
          "post_date": "06/19/2021 02:00:02",
          "content": "<p>thx for your reply!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1356382,
      "author_name": "mlconsult",
      "author_url": "",
      "post_date": "06/19/2021 01:42:15",
      "content": "<p>I recommend looking at these two notebooks and building on similar features. Good luck!  Make sure you up-vote any code you reference/use so they get credit.</p>\n<p><a href=\"https://www.kaggle.com/ulrich07/mlb-ann-with-lags-tf-keras\" target=\"_blank\">https://www.kaggle.com/ulrich07/mlb-ann-with-lags-tf-keras</a></p>\n<p><a href=\"https://www.kaggle.com/columbia2131/mlb-lightgbm-starter-dataset-code-en-ja\" target=\"_blank\">https://www.kaggle.com/columbia2131/mlb-lightgbm-starter-dataset-code-en-ja</a></p>\n<p><a href=\"https://www.kaggle.com/ulrich07\" target=\"_blank\">@ulrich07</a> <a href=\"https://www.kaggle.com/columbia2131\" target=\"_blank\">@columbia2131</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1356392,
          "author_name": "vanle73",
          "author_url": "",
          "post_date": "06/19/2021 02:00:43",
          "content": "<p>thx, this will be a great help!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1392472,
      "author_name": "eneszvo",
      "author_url": "",
      "post_date": "07/18/2021 18:13:37",
      "content": "<p>Give a try with <code>tsfresh</code> package if you need time-series features <a href=\"https://tsfresh.readthedocs.io/en/latest/\" target=\"_blank\">https://tsfresh.readthedocs.io/en/latest/</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1351539": "It will be grateful for providing some efficient pipelines to start a feature engineering competition. 🤓",
    "1354566": "There are all kinds of levels to feature engineering from the most basic to really advanced. I'm not sure what you are looking for exactly.",
    "1356300": "Most of feature engineering is basically common sense and creativity. Come up with ideas first, then figure out how to code them. Many ideas can be done with simple pandas functionality like group by and merge. Others might be trickier coding problems. I'd start with basic 1-1 transformations like date -> year, month, day-of-month and day-of-week. Then I'd do basic summary stuff on all the corresponding tables, and merge those to your training set. For example, a flag for \"did they play yesterday\" (left join games data on playerid, and then use .isna()). Another example is Players' teams' number of followers (join roster to get team id, then join teamFollowers to that). Finally, I'd try some simple time-dependent features, like basically any other feature or target from a day earlier. One way to do this is to copy your features/targets table, replace the date with (date + 1), rename columns \"..._lagged\" and merge that back onto your training data set.\n\nGood luck!",
    "1356382": "I recommend looking at these two notebooks and building on similar features. Good luck!  Make sure you up-vote any code you reference/use so they get credit.\n\nhttps://www.kaggle.com/ulrich07/mlb-ann-with-lags-tf-keras\n\nhttps://www.kaggle.com/columbia2131/mlb-lightgbm-starter-dataset-code-en-ja\n\n@ulrich07 @columbia2131",
    "1356390": "thx for your reply!",
    "1356392": "thx, this will be a great help!",
    "1392472": "Give a try with `tsfresh` package if you need time-series features https://tsfresh.readthedocs.io/en/latest/"
  },
  "source": "meta"
}