{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# MLB Feature Importance with LOFO\n\n![](https://upload.wikimedia.org/wikipedia/tr/4/48/MLB_Belirtke.png)\n![](https://raw.githubusercontent.com/aerdem4/lofo-importance/master/docs/lofo_logo.png)\n\n**LOFO** (Leave One Feature Out) Importance calculates the importances of a set of features based on **a metric of choice**, for **a model of choice**, by **iteratively removing each feature from the set**, and **evaluating the performance** of the model, with **a validation scheme of choice**, based on the chosen metric.\n\nLOFO first evaluates the performance of the model with all the input features included, then iteratively removes one feature at a time, retrains the model, and evaluates its performance on a validation set. The mean and standard deviation (across the folds) of the importance of each feature is then reported.\n\nWhile other feature importance methods usually calculate how much a feature is used by the model, LOFO estimates how much a feature can make a difference by itself given that we have the other features. Here are some advantages of LOFO:\n* It generalises well to unseen test sets since it uses a validation scheme.\n* It is model agnostic.\n* It gives negative importance to features that hurt performance upon inclusion.\n* It can group the features. Especially useful for high dimensional features like TFIDF or OHE features. It is also good practice to group very correlated features to avoid misleading results.\n\nhttps://github.com/aerdem4/lofo-importance","metadata":{}},{"cell_type":"code","source":"!pip install lofo-importance","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport os, gc\nfrom tqdm import tqdm\n\n\nROOT_DIR = \"/kaggle/input/mlb-player-digital-engagement-forecasting\"\n\ndf = pd.read_csv(f\"{ROOT_DIR}/train_updated.csv\")\nprint(df.shape)\ndf.head()","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","execution":{"iopub.status.busy":"2021-07-29T12:41:34.778652Z","iopub.execute_input":"2021-07-29T12:41:34.779023Z","iopub.status.idle":"2021-07-29T12:42:58.947176Z","shell.execute_reply.started":"2021-07-29T12:41:34.77899Z","shell.execute_reply":"2021-07-29T12:42:58.946193Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Read the data","metadata":{}},{"cell_type":"code","source":"target_df = [eval(x) for x in tqdm(df[\"nextDayPlayerEngagement\"].values)]\n\nflatten = lambda t: [item for sublist in t for item in sublist]\n\n\ntarget_df = pd.DataFrame(flatten(target_df))\n\nprint(target_df.shape)\ntarget_df.head()","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:42:58.948581Z","iopub.execute_input":"2021-07-29T12:42:58.948872Z","iopub.status.idle":"2021-07-29T12:44:22.072243Z","shell.execute_reply.started":"2021-07-29T12:42:58.948844Z","shell.execute_reply":"2021-07-29T12:44:22.071042Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"roster_df = [eval(x) for x in tqdm(df[\"rosters\"].values) if str(x) != \"nan\"]\n\nroster_df = pd.DataFrame(flatten(roster_df)).drop([\"statusCode\"], axis=1).rename(columns={\"gameDate\": \"date\"})\n\nprint(roster_df.shape)\nroster_df.head()","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:44:22.074967Z","iopub.execute_input":"2021-07-29T12:44:22.075319Z","iopub.status.idle":"2021-07-29T12:44:57.046967Z","shell.execute_reply.started":"2021-07-29T12:44:22.075279Z","shell.execute_reply":"2021-07-29T12:44:57.045947Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"standings_df = [eval(x.replace(\":false\", \":False\").replace(\":true\", \":True\").replace(\":null\", \":None\")) for x in tqdm(df[\"standings\"].values) if str(x) != \"nan\"]\n\nstandings_df = pd.DataFrame(flatten(standings_df)).rename(columns={\"gameDate\": \"date\"})[[\"date\", \"teamId\", \"leagueRank\", \"lastTenWins\"]]\n\nprint(standings_df.shape)\nstandings_df.head()","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:44:57.049142Z","iopub.execute_input":"2021-07-29T12:44:57.049593Z","iopub.status.idle":"2021-07-29T12:44:59.849021Z","shell.execute_reply.started":"2021-07-29T12:44:57.049548Z","shell.execute_reply":"2021-07-29T12:44:59.848105Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"transaction_df = []\n\nfor x in tqdm(df[\"transactions\"].values):\n    if str(x) != \"nan\":\n        transaction_df.extend(eval(x.replace(\":null\", ':\"\"')))\n\ntransaction_df = pd.DataFrame(transaction_df)\n\nprint(transaction_df.shape)\ntransaction_df.head()","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:44:59.850469Z","iopub.execute_input":"2021-07-29T12:44:59.850764Z","iopub.status.idle":"2021-07-29T12:45:01.859784Z","shell.execute_reply.started":"2021-07-29T12:44:59.850737Z","shell.execute_reply":"2021-07-29T12:45:01.858605Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"scores_df = []\n\nfor x in tqdm(df[\"playerBoxScores\"].values):\n    if str(x) != \"nan\":\n        scores_df.extend(eval(x.replace(\":null\", ':\"\"')))\n\nscores_df = pd.DataFrame(scores_df)\n\nprint(scores_df.shape)\nscores_df.head()","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:45:01.861215Z","iopub.execute_input":"2021-07-29T12:45:01.86153Z","iopub.status.idle":"2021-07-29T12:46:08.000979Z","shell.execute_reply.started":"2021-07-29T12:45:01.861489Z","shell.execute_reply":"2021-07-29T12:46:07.999975Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"awards_df = [eval(x.replace(\":null\", ':\"\"')) for x in df[\"awards\"].values if str(x) != \"nan\"]\nawards_df = pd.DataFrame(flatten(awards_df))\nawards_df.shape","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:46:08.002468Z","iopub.execute_input":"2021-07-29T12:46:08.002827Z","iopub.status.idle":"2021-07-29T12:46:08.174796Z","shell.execute_reply.started":"2021-07-29T12:46:08.002794Z","shell.execute_reply":"2021-07-29T12:46:08.173684Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"twitter_df = []\n\nfor x in tqdm(df[\"playerTwitterFollowers\"].values):\n    if str(x) != \"nan\":\n        twitter_df.extend(eval(x.replace(\":null\", ':\"\"')))\n\ntwitter_df = pd.DataFrame(twitter_df)[[\"date\", \"playerId\", \"numberOfFollowers\"]]\ntwitter_df[\"date\"] = pd.to_datetime(twitter_df[\"date\"])\n\nprint(twitter_df.shape)\ntwitter_df.head()","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:46:08.178389Z","iopub.execute_input":"2021-07-29T12:46:08.178701Z","iopub.status.idle":"2021-07-29T12:46:09.412701Z","shell.execute_reply.started":"2021-07-29T12:46:08.178672Z","shell.execute_reply":"2021-07-29T12:46:09.411662Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Only the evaluated players will be used for this analysis","metadata":{}},{"cell_type":"code","source":"players_df = pd.read_csv(f\"{ROOT_DIR}/players.csv\")\n\navailable_players = players_df[players_df[\"playerForTestSetAndFuturePreds\"] == True][\"playerId\"].values\nlen(available_players)","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:46:09.415175Z","iopub.execute_input":"2021-07-29T12:46:09.415608Z","iopub.status.idle":"2021-07-29T12:46:09.446921Z","shell.execute_reply.started":"2021-07-29T12:46:09.415566Z","shell.execute_reply":"2021-07-29T12:46:09.445636Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"target_df = target_df[target_df[\"playerId\"].isin(available_players)].reset_index(drop=True)\ntarget_df.shape","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:46:09.448618Z","iopub.execute_input":"2021-07-29T12:46:09.44903Z","iopub.status.idle":"2021-07-29T12:46:09.610826Z","shell.execute_reply.started":"2021-07-29T12:46:09.448988Z","shell.execute_reply":"2021-07-29T12:46:09.609969Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"targets = [\"target1\", \"target2\", \"target3\", \"target4\"]\n\n\ntarget_df[\"target_date\"] = pd.to_datetime(target_df[\"engagementMetricsDate\"])\ntarget_df['date'] = target_df['target_date'] -  pd.to_timedelta(1, unit='d')\n\ntarget_df[\"dom\"] = target_df[\"target_date\"].dt.day\ntarget_df[\"dow\"] = target_df[\"target_date\"].dt.dayofweek\ntarget_df[\"month\"] = target_df[\"target_date\"].dt.month - 1\ntarget_df[\"year\"] = target_df[\"target_date\"].dt.year\n\ntarget_df[\"time\"] = target_df[\"year\"]*12 + target_df[\"month\"]","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:46:09.612716Z","iopub.execute_input":"2021-07-29T12:46:09.613157Z","iopub.status.idle":"2021-07-29T12:46:10.522823Z","shell.execute_reply.started":"2021-07-29T12:46:09.613113Z","shell.execute_reply":"2021-07-29T12:46:10.521861Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Player Feature Extraction\n* Age of player\n* If the player is from US or not\n* Time since his debut date\n* Is he a jr (son of a previous well-known MLB player) ?","metadata":{}},{"cell_type":"code","source":"target_df = target_df.merge(players_df[[\"playerId\", \"DOB\", \"mlbDebutDate\", \"birthCountry\", \"playerName\"]],\n                            on=\"playerId\", how=\"left\") ","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:46:10.524444Z","iopub.execute_input":"2021-07-29T12:46:10.524906Z","iopub.status.idle":"2021-07-29T12:46:10.911348Z","shell.execute_reply.started":"2021-07-29T12:46:10.524864Z","shell.execute_reply":"2021-07-29T12:46:10.910441Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"target_df[\"age\"] = (target_df[\"date\"] - pd.to_datetime(target_df[\"DOB\"])).dt.days\ntarget_df[\"US_person\"] = 1*(target_df[\"birthCountry\"] == \"USA\")\ntarget_df[\"time_since_debut\"] = (target_df[\"date\"] - pd.to_datetime(target_df[\"mlbDebutDate\"])).dt.days\ntarget_df[\"jr\"] = target_df[\"playerName\"].apply(lambda x: \"Jr.\" in x)\ntarget_df[\"jr\"].mean()","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:46:10.9126Z","iopub.execute_input":"2021-07-29T12:46:10.9129Z","iopub.status.idle":"2021-07-29T12:46:12.578023Z","shell.execute_reply.started":"2021-07-29T12:46:10.912873Z","shell.execute_reply":"2021-07-29T12:46:12.577005Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Twitter Feature Extraction\n* Number of followers\n* Twitter trend: Percentage increase/decrease in followers compared to last month","metadata":{}},{"cell_type":"code","source":"twitter_df[\"twitter_trend\"] = ((twitter_df[\"numberOfFollowers\"] + 1 - twitter_df.groupby(\"playerId\")[\"numberOfFollowers\"].shift())\n                               / (twitter_df[\"numberOfFollowers\"] + 1))\ntwitter_df[\"twitter_trend\"].clip(-0.2, 0.2).hist(bins=50)","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:46:12.579316Z","iopub.execute_input":"2021-07-29T12:46:12.579632Z","iopub.status.idle":"2021-07-29T12:46:12.950252Z","shell.execute_reply.started":"2021-07-29T12:46:12.5796Z","shell.execute_reply":"2021-07-29T12:46:12.949348Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"target_df = target_df.merge(twitter_df, on=[\"date\", \"playerId\"], how=\"left\")\n\ntarget_df[\"numberOfFollowers\"] = target_df.groupby(\"playerId\")[\"numberOfFollowers\"].fillna(method=\"ffill\")\ntarget_df[\"twitter_trend\"] = target_df.groupby(\"playerId\")[\"twitter_trend\"].fillna(method=\"ffill\")","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:46:12.951818Z","iopub.execute_input":"2021-07-29T12:46:12.952234Z","iopub.status.idle":"2021-07-29T12:46:15.709144Z","shell.execute_reply.started":"2021-07-29T12:46:12.95219Z","shell.execute_reply":"2021-07-29T12:46:15.708177Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Awards Feature Extraction\n* Award type\n* Time since award","metadata":{}},{"cell_type":"code","source":"awards_df.loc[awards_df.groupby(\"awardId\")[\"awardDate\"].transform(\"count\") < 3, \"awardName\"] = \"Other\"\nawards_df[\"awardName\"].value_counts()","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:46:15.711014Z","iopub.execute_input":"2021-07-29T12:46:15.711478Z","iopub.status.idle":"2021-07-29T12:46:15.731063Z","shell.execute_reply.started":"2021-07-29T12:46:15.711431Z","shell.execute_reply":"2021-07-29T12:46:15.72999Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.preprocessing import LabelEncoder\n\nle = LabelEncoder()\nawards_df[\"awardType\"] = le.fit_transform(awards_df[\"awardName\"])","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:46:15.734155Z","iopub.execute_input":"2021-07-29T12:46:15.734443Z","iopub.status.idle":"2021-07-29T12:46:16.686791Z","shell.execute_reply.started":"2021-07-29T12:46:15.734415Z","shell.execute_reply":"2021-07-29T12:46:16.684775Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"awards_df = awards_df.groupby([\"awardDate\", \"playerId\"])[\"awardType\"].first().reset_index().rename(columns={\"awardDate\": \"date\"})\nawards_df[\"date\"] = pd.to_datetime(awards_df[\"date\"])\nawards_df[\"award_date\"] = awards_df[\"date\"].values\nawards_df.head()","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:46:16.68802Z","iopub.execute_input":"2021-07-29T12:46:16.688318Z","iopub.status.idle":"2021-07-29T12:46:16.72073Z","shell.execute_reply.started":"2021-07-29T12:46:16.68829Z","shell.execute_reply":"2021-07-29T12:46:16.719732Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"target_df = target_df.merge(awards_df, how=\"left\", on=[\"date\", \"playerId\"])\ntarget_df[\"award_date\"].isnull().mean()","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:46:16.722064Z","iopub.execute_input":"2021-07-29T12:46:16.722379Z","iopub.status.idle":"2021-07-29T12:46:17.37576Z","shell.execute_reply.started":"2021-07-29T12:46:16.72235Z","shell.execute_reply":"2021-07-29T12:46:17.37461Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"target_df[\"award_date\"] = target_df.groupby(\"playerId\")[\"award_date\"].fillna(method=\"ffill\")\ntarget_df[\"awardType\"] = target_df.groupby(\"playerId\")[\"awardType\"].fillna(method=\"ffill\")","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:46:17.377368Z","iopub.execute_input":"2021-07-29T12:46:17.377722Z","iopub.status.idle":"2021-07-29T12:46:19.585761Z","shell.execute_reply.started":"2021-07-29T12:46:17.377691Z","shell.execute_reply":"2021-07-29T12:46:19.584754Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"target_df[\"time_since_award\"] = (target_df[\"date\"] - target_df[\"award_date\"]).dt.days\ntarget_df[\"time_since_award\"].hist()","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:46:19.587434Z","iopub.execute_input":"2021-07-29T12:46:19.588165Z","iopub.status.idle":"2021-07-29T12:46:19.912738Z","shell.execute_reply.started":"2021-07-29T12:46:19.5881Z","shell.execute_reply":"2021-07-29T12:46:19.911718Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Roster Feature Extraction\n* Status","metadata":{}},{"cell_type":"code","source":"roster_df.loc[roster_df.groupby(\"status\")[\"date\"].transform(\"count\") < 20, \"status\"] = \"Other\"\nroster_df[\"status\"].value_counts()","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:46:19.914154Z","iopub.execute_input":"2021-07-29T12:46:19.914855Z","iopub.status.idle":"2021-07-29T12:46:20.604147Z","shell.execute_reply.started":"2021-07-29T12:46:19.914804Z","shell.execute_reply":"2021-07-29T12:46:20.603066Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"le = LabelEncoder()\nroster_df[\"status\"] = le.fit_transform(roster_df[\"status\"])","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:46:20.607951Z","iopub.execute_input":"2021-07-29T12:46:20.60826Z","iopub.status.idle":"2021-07-29T12:46:21.131891Z","shell.execute_reply.started":"2021-07-29T12:46:20.608231Z","shell.execute_reply":"2021-07-29T12:46:21.130817Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"roster_df[\"date\"] = pd.to_datetime(roster_df[\"date\"])\n\ntarget_df = target_df.merge(roster_df, how=\"left\", on=[\"date\", \"playerId\"])\ntarget_df[\"status\"].isnull().mean(), target_df[\"teamId\"].isnull().mean()","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:46:21.133822Z","iopub.execute_input":"2021-07-29T12:46:21.13414Z","iopub.status.idle":"2021-07-29T12:46:22.456317Z","shell.execute_reply.started":"2021-07-29T12:46:21.13411Z","shell.execute_reply":"2021-07-29T12:46:22.455342Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Transaction Feature Extraction\n* Transfer type\n* Time since transfer","metadata":{}},{"cell_type":"code","source":"transaction_df[\"date\"] = pd.to_datetime(transaction_df[\"date\"])\ntransaction_df[\"effectiveDate\"] = pd.to_datetime(transaction_df[\"effectiveDate\"])\n\n\nle = LabelEncoder()\ntransaction_df[\"transfer_type\"] = le.fit_transform(transaction_df[\"typeCode\"])","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:46:22.457493Z","iopub.execute_input":"2021-07-29T12:46:22.457787Z","iopub.status.idle":"2021-07-29T12:46:22.513163Z","shell.execute_reply.started":"2021-07-29T12:46:22.45776Z","shell.execute_reply":"2021-07-29T12:46:22.512039Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"target_df = target_df.merge(transaction_df[[\"playerId\", \"date\", \"effectiveDate\", \"transfer_type\"]], \n                            how=\"left\", on=[\"date\", \"playerId\"])","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:46:22.514471Z","iopub.execute_input":"2021-07-29T12:46:22.514863Z","iopub.status.idle":"2021-07-29T12:46:23.802948Z","shell.execute_reply.started":"2021-07-29T12:46:22.514833Z","shell.execute_reply":"2021-07-29T12:46:23.801892Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"target_df[\"effectiveDate\"] = target_df.groupby(\"playerId\")[\"effectiveDate\"].fillna(method=\"ffill\")\ntarget_df[\"transfer_type\"] = target_df.groupby(\"playerId\")[\"transfer_type\"].fillna(method=\"ffill\")\n\ntarget_df[\"time_since_transfer\"] = (target_df[\"date\"] - target_df[\"effectiveDate\"]).dt.days\ntarget_df[\"time_since_transfer\"].hist()","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:46:23.804258Z","iopub.execute_input":"2021-07-29T12:46:23.804538Z","iopub.status.idle":"2021-07-29T12:46:26.378676Z","shell.execute_reply.started":"2021-07-29T12:46:23.804499Z","shell.execute_reply":"2021-07-29T12:46:26.377716Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Scores Feature Extraction\n* Total Bases\n* Strike outs Pitching\n* Plate Appearances\n* Home Runs\n* Innings Pitched\n* Saves\n* rbi\n* No game since","metadata":{}},{"cell_type":"code","source":"scores_df.rename(columns={\"gameDate\": \"date\"}, inplace=True)\nscores_df[\"date\"] = pd.to_datetime(scores_df[\"date\"])\n\nscores_features = ['totalBases', 'strikeOutsPitching', 'plateAppearances', \n                   'homeRuns', 'inningsPitched', 'saves', 'rbi']\ntarget_df = target_df.merge(scores_df[[\"playerId\", \"date\"] + scores_features], on=[\"playerId\", \"date\"], how=\"left\")\n\ntarget_df[scores_features].isnull().mean()","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:46:26.379902Z","iopub.execute_input":"2021-07-29T12:46:26.380186Z","iopub.status.idle":"2021-07-29T12:46:29.071668Z","shell.execute_reply.started":"2021-07-29T12:46:26.380159Z","shell.execute_reply":"2021-07-29T12:46:29.070793Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def f(x):\n    y = np.zeros(len(x))\n    \n    for i, el in enumerate(x):\n        if el or (i == 0):\n            y[i] = 0\n        else:\n            y[i] = y[i-1] + 1\n    return y\n            \n\ntarget_df[\"no_game_since\"] = target_df[\"rbi\"].notnull()\ntarget_df[\"no_game_since\"] = target_df.groupby(\"playerId\")[\"no_game_since\"].transform(f)\ntarget_df[\"no_game_since\"].hist()","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:46:29.072909Z","iopub.execute_input":"2021-07-29T12:46:29.073182Z","iopub.status.idle":"2021-07-29T12:46:32.025449Z","shell.execute_reply.started":"2021-07-29T12:46:29.073157Z","shell.execute_reply":"2021-07-29T12:46:32.024611Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in tqdm(scores_features):\n    target_df[col] = target_df.groupby(\"playerId\")[col].fillna(method=\"ffill\")","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:46:32.026958Z","iopub.execute_input":"2021-07-29T12:46:32.027407Z","iopub.status.idle":"2021-07-29T12:46:46.972069Z","shell.execute_reply.started":"2021-07-29T12:46:32.027362Z","shell.execute_reply":"2021-07-29T12:46:46.970872Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def to_float(x):\n    try:\n        return float(x)\n    except:\n        pass\n    return None\n\n\nfor col in scores_features:\n    target_df[col] = target_df[col].apply(to_float)\n    \ntarget_df[\"playerId\"] = target_df[\"playerId\"].astype(int)","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:46:46.973648Z","iopub.execute_input":"2021-07-29T12:46:46.974068Z","iopub.status.idle":"2021-07-29T12:46:55.417795Z","shell.execute_reply.started":"2021-07-29T12:46:46.974025Z","shell.execute_reply":"2021-07-29T12:46:55.416849Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Standing Features\n* League Rank\n* Last Ten Wins","metadata":{}},{"cell_type":"code","source":"standings_df[\"date\"] = pd.to_datetime(standings_df[\"date\"])\n\ntarget_df = target_df.merge(standings_df, on=[\"teamId\", \"date\"], how=\"left\")\nstanding_features = [\"leagueRank\", \"lastTenWins\"]\n\nfor col in standing_features:\n    target_df[col] = target_df.groupby(\"playerId\")[col].fillna(method=\"ffill\")","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:46:55.419406Z","iopub.execute_input":"2021-07-29T12:46:55.419824Z","iopub.status.idle":"2021-07-29T12:46:59.544306Z","shell.execute_reply.started":"2021-07-29T12:46:55.419781Z","shell.execute_reply":"2021-07-29T12:46:59.54357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Target Lag Features\n\n**Since the competition task is to predict the digital engagement for 45 days, we need to be careful with lag features. We won't have yesterday's target values for day 10 for example. In order to simplify the problem for feature importance, we can assume that we want to predict for the 20th day. So all lag features should be calculated before the 20th day. Mean and std for the last available month and year are calculated.**","metadata":{}},{"cell_type":"code","source":"FUTURE = 20\n\nlags = np.array([21, 28, 35])\nassert np.all(lags > FUTURE)\n\n\n\nfor lag in lags:\n    for t in tqdm(targets):\n        fname = f\"lag{lag}_{t}\"\n        target_df[fname] = target_df.groupby(\"playerId\")[t].shift(lag)\n        \n        if lag == 28:\n            for period in [4*7, 52*7]:\n                fname = f\"std{period}_{lag}_{t}\"\n                target_df[fname] = target_df.groupby(\"playerId\")[t].rolling(period).std().reset_index().sort_values(\"level_1\")[t].values\n                target_df[fname] = target_df.groupby(\"playerId\")[fname].shift(lag)\n                fname = f\"mean{period}_{lag}_{t}\"\n                target_df[fname] = target_df.groupby(\"playerId\")[t].rolling(period).mean().reset_index().sort_values(\"level_1\")[t].values\n                target_df[fname] = target_df.groupby(\"playerId\")[fname].shift(lag)","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:46:59.545716Z","iopub.execute_input":"2021-07-29T12:46:59.546304Z","iopub.status.idle":"2021-07-29T12:47:23.913009Z","shell.execute_reply.started":"2021-07-29T12:46:59.546259Z","shell.execute_reply":"2021-07-29T12:47:23.912276Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for t in targets:\n    target_df[f\"dif_{t}\"] = target_df[f\"lag21_{t}\"] - target_df[f\"lag28_{t}\"]\n    target_df[f\"diff_{t}\"] = target_df[f\"dif_{t}\"] - (target_df[f\"lag28_{t}\"] - target_df[f\"lag35_{t}\"])","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:47:23.914211Z","iopub.execute_input":"2021-07-29T12:47:23.914702Z","iopub.status.idle":"2021-07-29T12:47:23.974252Z","shell.execute_reply.started":"2021-07-29T12:47:23.91467Z","shell.execute_reply":"2021-07-29T12:47:23.973214Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Validation Scheme: Time-based Cross-validation\n\n**We need to split the data by time, because our model gets applied on future data. But one time split is not enough, therefore we can pick 5 months as validation sets and the other months before them as training set. The competition will run on August-September, therefore these 2 months are good choices to be included in validation and at least one month from 2021 would be nice.**\n* August 2019\n* September 2019\n* August 2020\n* September 2020\n* May 2021\n\n![](https://miro.medium.com/max/558/1*AXRu72CV1hdjLfODFGbMWQ.png)","metadata":{}},{"cell_type":"code","source":"VAL_MONTHS = {12*2019 + 7, 12*2019 + 8, 12*2020 + 7, 12*2020 + 8, 12*2021 + 4}\n# 2019 August, September, 2020 August, September, 2021 May\nVAL_MONTHS = target_df[(target_df[\"year\"]*12 + target_df[\"month\"]).isin(VAL_MONTHS)][\"time\"].drop_duplicates().values\n\ncv_scheme = []\n\nfor valm in VAL_MONTHS:\n    train_ind = np.where(target_df[\"time\"].values < valm)[0]\n    val_ind = np.where(target_df[\"time\"].values == valm)[0]\n    cv_scheme.append((train_ind, val_ind))\nVAL_MONTHS","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:47:23.975539Z","iopub.execute_input":"2021-07-29T12:47:23.975849Z","iopub.status.idle":"2021-07-29T12:47:25.440041Z","shell.execute_reply.started":"2021-07-29T12:47:23.975819Z","shell.execute_reply":"2021-07-29T12:47:25.439077Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Individual features including categoricals","metadata":{}},{"cell_type":"code","source":"categoricals = [\"year\", \"dow\", \"awardType\", \"transfer_type\", \"status\", \"playerId\", \"teamId\", \"month\"]\nnumericals = [\"time_since_award\", \"time_since_transfer\", \"no_game_since\",\n              \"age\", \"time_since_debut\", \"US_person\", \"jr\", \"twitter_trend\", \"numberOfFollowers\"]\nfeatures = list(categoricals) + list(numericals)","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:47:25.44272Z","iopub.execute_input":"2021-07-29T12:47:25.443185Z","iopub.status.idle":"2021-07-29T12:47:25.448183Z","shell.execute_reply.started":"2021-07-29T12:47:25.443139Z","shell.execute_reply":"2021-07-29T12:47:25.447172Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### LOFO allows us to group features. This prevents highly correlated features to have underestimated importance.","metadata":{}},{"cell_type":"code","source":"target_features = {f\"{t}_features\": target_df[[f\"lag21_{t}\", f\"dif_{t}\", f\"diff_{t}\",\n                                               f\"std28_28_{t}\", f\"mean28_28_{t}\",\n                                               f\"std364_28_{t}\", f\"mean364_28_{t}\"]].values for t in targets}\n\ntarget_features[\"scores_features\"] = target_df[scores_features].values\ntarget_features[\"standing_features\"] = target_df[standing_features].values\n\ntarget_features.keys()","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:47:25.450059Z","iopub.execute_input":"2021-07-29T12:47:25.450785Z","iopub.status.idle":"2021-07-29T12:47:25.643852Z","shell.execute_reply.started":"2021-07-29T12:47:25.450742Z","shell.execute_reply":"2021-07-29T12:47:25.642704Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Applying LOFO\n* Model: LGBMRegressor with Huber loss\n* Metric: Mean Absolute Error\n* CV: 5 time splits\n\n**Green bars represent useful, red bars represent harmful features. We can also see the standard deviation of importance. This helps us understand if the feature is really important or it is a noisy estimation.**","metadata":{}},{"cell_type":"markdown","source":"# target1","metadata":{}},{"cell_type":"code","source":"from lofo import Dataset, LOFOImportance, plot_importance\nfrom lightgbm import LGBMClassifier, LGBMRegressor\n\n\ndef get_importance(target_name):\n    model = LGBMRegressor(min_child_samples=20, n_jobs=-1, objective=\"huber\", alpha=0.1,\n                          n_estimators=100)\n    dataset = Dataset(df=target_df, target=target_name, features=features, feature_groups=target_features)\n    lofo_imp = LOFOImportance(dataset, cv=cv_scheme, scoring=\"neg_mean_absolute_error\", model=model,\n                              fit_params={\"categorical_feature\": categoricals})\n    return lofo_imp.get_importance()\n\n\nimportance_df = get_importance(\"target1\")\nplot_importance(importance_df, figsize=(8, 8), kind=\"default\")","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:49:41.906236Z","iopub.execute_input":"2021-07-29T12:49:41.906626Z","iopub.status.idle":"2021-07-29T12:55:08.263493Z","shell.execute_reply.started":"2021-07-29T12:49:41.906588Z","shell.execute_reply":"2021-07-29T12:55:08.258831Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**scores_features** and **no_games_since** are the only significantly important features. teamId, target1_features, status, month also look important but not significant enough. **playerId** seems to be a harmful feature for target1, meaning that it causes overfitting and we can benefit from removing this feature. Most of the features seem to be either redundant or not useful. One would expect lag features of **target1** to have the highest importance but LOFO Importance is the importance of a feature given that you have the other features. So it is about how much it can make a difference. Having target2, target3, target4 lag features and playerId, target1 lag features seem to be redundant.\n\nLet's repeat the experiment with less features:\n","metadata":{}},{"cell_type":"code","source":"model = LGBMRegressor(min_child_samples=20, n_jobs=-1, objective=\"huber\", alpha=0.1,\n                      n_estimators=100)\ndataset = Dataset(df=target_df, target=\"target1\", \n                  features=[\"no_game_since\", \"teamId\", \"status\", \"month\"], \n                  feature_groups={\"target1_features\": target_features[\"target1_features\"],\n                                  \"scores_features\": target_features[\"scores_features\"]})\nlofo_imp = LOFOImportance(dataset, cv=cv_scheme, scoring=\"neg_mean_absolute_error\", model=model,\n                          fit_params={\"categorical_feature\": [\"teamId\", \"status\", \"month\"]})\n\nplot_importance(lofo_imp.get_importance(), figsize=(8, 4), kind=\"box\")","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:55:08.264819Z","iopub.status.idle":"2021-07-29T12:55:08.265265Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**target1** lag features now have larger importance but still less than **scores_features** and **no_games_since**. We can say that the target values for previous month and year is not very important for **target1**. **target1** mainly depends on **Player Box Scores** data.","metadata":{}},{"cell_type":"markdown","source":"# target2","metadata":{}},{"cell_type":"code","source":"importance_df = get_importance(\"target2\")\nplot_importance(importance_df, figsize=(8, 8), kind=\"default\")","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:55:08.26655Z","iopub.status.idle":"2021-07-29T12:55:08.267002Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**status**, **target2_features**, **scores_features**, **no_game_since** and **month** seem to be the important features. Let's repeat the experiment with less features.","metadata":{}},{"cell_type":"code","source":"model = LGBMRegressor(min_child_samples=20, n_jobs=-1, objective=\"huber\", alpha=0.1,\n                      n_estimators=100)\ndataset = Dataset(df=target_df, target=\"target2\", \n                  features=[\"no_game_since\", \"teamId\", \"status\", \"month\"], \n                  feature_groups={\"target2_features\": target_features[\"target2_features\"],\n                                  \"scores_features\": target_features[\"scores_features\"]})\nlofo_imp = LOFOImportance(dataset, cv=cv_scheme, scoring=\"neg_mean_absolute_error\", model=model,\n                          fit_params={\"categorical_feature\": [\"teamId\", \"status\", \"month\"]})\n\nplot_importance(lofo_imp.get_importance(), figsize=(8, 4), kind=\"box\")","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:55:08.267898Z","iopub.status.idle":"2021-07-29T12:55:08.268319Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**no_game_since** and **scores_features** are important but not as much as they were for **target1**. And **month** feature has high std meaning that it is probably not consistently improtant across the all validation sets. The most important features are **target2** lag features and **status** feature.","metadata":{}},{"cell_type":"markdown","source":"# target3","metadata":{}},{"cell_type":"code","source":"importance_df = get_importance(\"target3\")\nplot_importance(importance_df, figsize=(8, 8), kind=\"default\")","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:55:08.269189Z","iopub.status.idle":"2021-07-29T12:55:08.269624Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**scores_features** and **all target lag features** seem to be important and **no_game_since** is a harmful feature. We can repeat the experiment with less features:","metadata":{}},{"cell_type":"code","source":"model = LGBMRegressor(min_child_samples=20, n_jobs=-1, objective=\"huber\", alpha=0.1,\n                      n_estimators=100)\ndataset = Dataset(df=target_df, target=\"target3\", \n                  features=[\"no_game_since\", \"time_since_transfer\", \"status\", \"month\"], \n                  feature_groups={\"target3_features\": target_features[\"target3_features\"],\n                                  \"scores_features\": target_features[\"scores_features\"]})\nlofo_imp = LOFOImportance(dataset, cv=cv_scheme, scoring=\"neg_mean_absolute_error\", model=model,\n                          fit_params={\"categorical_feature\": [\"status\", \"month\"]})\n\nplot_importance(lofo_imp.get_importance(), figsize=(8, 4), kind=\"box\")","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:55:08.270644Z","iopub.status.idle":"2021-07-29T12:55:08.27105Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**target3** lag features are the most important ones. One interesting thing is that while **scores_features** are important, **no_game_since** feature is harmful. For some reason, **Player Box Scores** data is useful for **target3** but only the recent scores without knowing how many days old they are. But there is very small signal overall for **target**, it is difficult to make strong conclusions.","metadata":{}},{"cell_type":"markdown","source":"# target4","metadata":{}},{"cell_type":"code","source":"importance_df = get_importance(\"target4\")\nplot_importance(importance_df, figsize=(8, 8), kind=\"default\")","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:55:08.271949Z","iopub.status.idle":"2021-07-29T12:55:08.272371Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**target4** lag features, **playerId** and **status** are important features but they have small mean/std ratios. So they are not robust across different validation sets. **time_since_debut** is a consistently harmful feature but its performance degradation is not significant.","metadata":{}},{"cell_type":"code","source":"model = LGBMRegressor(min_child_samples=20, n_jobs=-1, objective=\"huber\", alpha=0.1,\n                      n_estimators=100)\ndataset = Dataset(df=target_df, target=\"target4\", \n                  features=[\"no_game_since\", \"playerId\", \"status\", \"month\"], \n                  feature_groups={\"target4_features\": target_features[\"target4_features\"],\n                                  \"scores_features\": target_features[\"scores_features\"]})\nlofo_imp = LOFOImportance(dataset, cv=cv_scheme, scoring=\"neg_mean_absolute_error\", model=model,\n                          fit_params={\"categorical_feature\": [\"playerId\", \"status\", \"month\"]})\n\nplot_importance(lofo_imp.get_importance(), figsize=(8, 4), kind=\"box\")","metadata":{"execution":{"iopub.status.busy":"2021-07-29T12:55:08.273164Z","iopub.status.idle":"2021-07-29T12:55:08.273627Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**no_game_since** is a harmful feature. It probably means that behavior of **no_game_since** changes over time. Learning its behavior in interaction with other features has more harm than its benefits for **target4**.","metadata":{}},{"cell_type":"markdown","source":"# Summary\n* **target1** mostly depends on the player's individual game performances.\n* Previous target2 values are very predictive for **target2** and status is also very important.\n* There is very little signal for **target3** and **target4**. Target lag features are useful and robust compared to other features. And **no_games_since** is a harmful feature for these targets.","metadata":{}}]}