{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<center><img src=\"https://stn2.tv/wp-content/uploads/2020/04/mlb-logo.jpg\"></center>","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport plotly.express as px\nimport seaborn as sns\nfrom tqdm import tqdm\nimport matplotlib.pyplot as plt\n#!pip install raceplotly\n#from raceplotly.plots import barplot","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2021-06-17T20:32:48.257456Z","iopub.execute_input":"2021-06-17T20:32:48.257892Z","iopub.status.idle":"2021-06-17T20:35:20.643421Z","shell.execute_reply.started":"2021-06-17T20:32:48.257797Z","shell.execute_reply":"2021-06-17T20:35:20.639339Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# About the Competition🚩\n<p style=\"font-size:15px\">In this competition, you’ll predict how fans engage with MLB players’ digital content on a daily basis for a future date range. You’ll have access to player performance data, social media data, and team factors like market size. Successful models will provide new insights into what signals most strongly correlate with and influence engagement.\n\nImagine if you could predict MLB All Stars all season long or when each of a team’s 25 players has his moment in the spotlight. These insights are possible when you dive deeper into the fandom of America’s pastime. Be part of the first method of its kind to try to understand digital engagement at the player level in this granular, day-to-day fashion. Simultaneously help MLB build innovation more easily using Google Cloud’s data analytics, Vertex AI and MLOps tools. You could play a part in shaping the future of MLB fan and player engagement.\n\nSubmissions are evaluated on the mean column-wise mean absolute error (MCMAE). A mean absolute error is calculated for each of the four target variables and the score is the average of those four MAE values.\n\n</p>","metadata":{}},{"cell_type":"markdown","source":"# Data Description","metadata":{}},{"cell_type":"markdown","source":"<div style=\"font-size:15px\">\n We are given 7 csv files:-\n<ul>\n    <li><code>train.csv:</code>training set</li>\n    <li><code>example_test.csv:</code>example of test set</li>\n    <li><code>example_sample_submission.csv:</code>example of sample_submission</li>\n    <li><code>awards.csv:</code>awards won by players before 2018</li>\n    <li><code>players.csv:</code>Library high level information about all players.</li>\n    <li><code>seasons.csv:</code>Information about start and end dates of all seasons in this dataset</li>\n    <li><code>teams.csv:</code>Library containing high level information about all MLB teams.</li>\n</ul>    \n</div>","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:15px; font-family:verdana; line-height: 2.0em;\">\nNote: Since this is a code competition You must submit to this competition using the provided MLB python time-series module, which ensures that models do not peek forward in time.\n</div>","metadata":{}},{"cell_type":"markdown","source":"# EDA","metadata":{}},{"cell_type":"code","source":"players = pd.read_csv('../input/mlb-player-digital-engagement-forecasting/players.csv')\nseasons = pd.read_csv('../input/mlb-player-digital-engagement-forecasting/seasons.csv')\nawards = pd.read_csv('../input/mlb-player-digital-engagement-forecasting/awards.csv')\nteams = pd.read_csv('../input/mlb-player-digital-engagement-forecasting/teams.csv')\ntrain = pd.read_csv('../input/mlb-player-digital-engagement-forecasting/train.csv')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-06-17T20:35:20.644762Z","iopub.status.idle":"2021-06-17T20:35:20.645124Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<p style=\"font-size:15px\">Let's take a peek at train.csv</p>","metadata":{}},{"cell_type":"code","source":"train.head()","metadata":{"execution":{"iopub.status.busy":"2021-06-17T20:35:20.645804Z","iopub.status.idle":"2021-06-17T20:35:20.646284Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.info()","metadata":{"execution":{"iopub.status.busy":"2021-06-17T20:35:20.647324Z","iopub.status.idle":"2021-06-17T20:35:20.647671Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:1ilocx; font-family:verdana; line-height: 2.0em;\">\nNote: Coulmn nextDayPlayerEngagement contains the targets that we want to predict\n</div>","metadata":{}},{"cell_type":"code","source":"train","metadata":{"execution":{"iopub.status.busy":"2021-06-17T20:35:20.648233Z","iopub.status.idle":"2021-06-17T20:35:20.648530Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I am using an unnested dataset created by @Ml_Bear link of the original [kernel](https://www.kaggle.com/naotaka1128/creating-unnested-dataset) do check it out","metadata":{}},{"cell_type":"code","source":"train_next_day = pd.read_pickle('../input/mlb-unnested/train_nextDayPlayerEngagement.pickle')\ntrain_next_day.engagementMetricsDate = train_next_day.engagementMetricsDate.astype('datetime64')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-06-17T20:35:20.649090Z","iopub.status.idle":"2021-06-17T20:35:20.649369Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_next_day","metadata":{"execution":{"iopub.status.busy":"2021-06-17T20:35:20.649979Z","iopub.status.idle":"2021-06-17T20:35:20.650280Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_next_day.head()","metadata":{"execution":{"iopub.status.busy":"2021-06-17T20:35:20.650916Z","iopub.status.idle":"2021-06-17T20:35:20.651224Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<p style=\"font-size:15px\">Let's visualize how our targets change over time</p>","metadata":{}},{"cell_type":"code","source":"unique_date = list(train_next_day.engagementMetricsDate.unique())\ntarget1_lis = []\ntarget2_lis = []\ntarget3_lis = []\ntarget4_lis = []\nfor i in unique_date:\n    df = train_next_day[train_next_day['engagementMetricsDate']==i]\n    target1 = (df['target1'].sum())/len(df)\n    target2 = (df['target2'].sum())/len(df)\n    target3 = (df['target3'].sum())/len(df)\n    target4 = (df['target4'].sum())/len(df)\n    target1_lis.append(target1)\n    target2_lis.append(target2)\n    target3_lis.append(target3)\n    target4_lis.append(target4)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-06-17T20:35:20.651972Z","iopub.status.idle":"2021-06-17T20:35:20.652255Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"px.line(x=unique_date,y=target1_lis,title='target 1 over time')","metadata":{"execution":{"iopub.status.busy":"2021-06-17T20:35:20.652778Z","iopub.status.idle":"2021-06-17T20:35:20.653106Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"px.line(x=unique_date,y=target2_lis,title='target 2 over time')","metadata":{"execution":{"iopub.status.busy":"2021-06-17T20:35:20.653867Z","iopub.status.idle":"2021-06-17T20:35:20.654173Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"px.line(x=unique_date,y=target3_lis,title='target 3 over time')","metadata":{"execution":{"iopub.status.busy":"2021-06-17T20:35:20.654792Z","iopub.status.idle":"2021-06-17T20:35:20.655094Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"px.line(x=unique_date,y=target4_lis,title='target 4 over time')","metadata":{"execution":{"iopub.status.busy":"2021-06-17T20:35:20.655736Z","iopub.status.idle":"2021-06-17T20:35:20.656041Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"team_twitter = pd.read_pickle('../input/mlb-unnested/train_teamTwitterFollowers.pickle')\nteam_twitter","metadata":{"execution":{"iopub.status.busy":"2021-06-17T20:35:20.657034Z","iopub.status.idle":"2021-06-17T20:35:20.657345Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<p style=\"font-size:15px\"> We can plot a raceplot to visualize no. of followers of the team growing over the years  </p>","metadata":{}},{"cell_type":"code","source":"#my_raceplot = barplot(team_twitter,  item_column='teamName', value_column='numberOfFollowers', time_column='date')\n#my_raceplot.plot(item_label = 'team name', value_label = 'number of followers', frame_duration = 800)","metadata":{"execution":{"iopub.status.busy":"2021-06-17T20:35:20.658242Z","iopub.status.idle":"2021-06-17T20:35:20.658597Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:1ilocx; font-family:verdana; line-height: 2.0em;\">\n📌no. of followers have grown over the years<br>\n📌Houston Astros also seems to becoming more popular\n</div>","metadata":{}},{"cell_type":"code","source":"standings = pd.read_pickle('../input/mlb-unnested/train_standings.pickle')\ntransactions = pd.read_pickle('../input/mlb-unnested/train_transactions.pickle')","metadata":{"execution":{"iopub.status.busy":"2021-06-17T20:35:20.659339Z","iopub.status.idle":"2021-06-17T20:35:20.659671Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"similary we can also plot raceplot to visuzalize no. of wins of a team over the years","metadata":{}},{"cell_type":"code","source":"#my_raceplot = barplot(standings,  item_column='teamName',value_column='wins', time_column='dailyDataDate',top_entries=10)\n#my_raceplot.plot(item_label = 'team name', value_label = 'current wins', frame_duration = 800)","metadata":{"execution":{"iopub.status.busy":"2021-06-17T20:35:20.660401Z","iopub.status.idle":"2021-06-17T20:35:20.660735Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"transactions = transactions.dropna()","metadata":{"execution":{"iopub.status.busy":"2021-06-17T20:35:20.661682Z","iopub.status.idle":"2021-06-17T20:35:20.662009Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"transactions","metadata":{"execution":{"iopub.status.busy":"2021-06-17T20:35:20.663876Z","iopub.status.idle":"2021-06-17T20:35:20.664327Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"most_transaction_team=transactions.groupby('fromTeamName')['toTeamName'].count().reset_index(name='Count').sort_values('Count',ascending=False)\npx.bar(most_transaction_team.head(10),x='fromTeamName',y='Count')","metadata":{"execution":{"iopub.status.busy":"2021-06-17T20:35:20.665058Z","iopub.status.idle":"2021-06-17T20:35:20.665404Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"most_transaction_player=transactions.groupby('playerName')['toTeamName'].count().reset_index(name='Count').sort_values('Count',ascending=False)\npx.bar(most_transaction_player.head(10),x='playerName',y='Count')","metadata":{"execution":{"iopub.status.busy":"2021-06-17T20:35:20.666169Z","iopub.status.idle":"2021-06-17T20:35:20.666535Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<p style=\"font-size:15px\">Now let's take look at players.csv</p>","metadata":{}},{"cell_type":"code","source":"players.head()","metadata":{"execution":{"iopub.status.busy":"2021-06-17T20:35:20.667249Z","iopub.status.idle":"2021-06-17T20:35:20.667636Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"players.info()","metadata":{"execution":{"iopub.status.busy":"2021-06-17T20:35:20.668393Z","iopub.status.idle":"2021-06-17T20:35:20.668738Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<p style=\"font-size:15px\">Let's see viz of country</p>","metadata":{}},{"cell_type":"code","source":"px.histogram(players,x='birthCountry',color='birthCountry')","metadata":{"execution":{"iopub.status.busy":"2021-06-17T20:35:20.669498Z","iopub.status.idle":"2021-06-17T20:35:20.669855Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"font-size:15px\">Let's see if there is a relationship between height and weight of players</div>","metadata":{}},{"cell_type":"code","source":"px.scatter(players,x='weight',y='heightInches')","metadata":{"execution":{"iopub.status.busy":"2021-06-17T20:35:20.670625Z","iopub.status.idle":"2021-06-17T20:35:20.670955Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"font-size:15px\">We can also combine 2 data frames to gain more insights for example by combining award count with primary position name we can see which position gets most awards</div>","metadata":{}},{"cell_type":"code","source":"playerid = list(awards['playerId'])\naward_count = []\nfor i in playerid:\n    award_count.append(len(awards[awards['playerId']==i]))\naward_count = pd.DataFrame({\"playerId\":playerid,\"award_count\":award_count})\nplayers = pd.merge(players,award_count,on='playerId')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-06-17T20:35:20.671684Z","iopub.status.idle":"2021-06-17T20:35:20.671993Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"postition_list = list(players['primaryPositionName'].unique())\naward_count_sum = []\nfor i in postition_list:\n    award_count_sum.append(players[players['primaryPositionName']==i]['award_count'].sum()) ","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-06-17T20:35:20.672744Z","iopub.status.idle":"2021-06-17T20:35:20.673055Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"px.bar(x=postition_list,y=award_count_sum)","metadata":{"execution":{"iopub.status.busy":"2021-06-17T20:35:20.673910Z","iopub.status.idle":"2021-06-17T20:35:20.674237Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"most_award = awards['playerId'].mode()\nprint(f\"playerID: {most_award.values}\")\nawards[awards['playerId'] == 405395]","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-06-17T20:35:20.674990Z","iopub.status.idle":"2021-06-17T20:35:20.675362Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"px.histogram(awards,y='awardName',category_orders=awards['awardName'])","metadata":{"execution":{"iopub.status.busy":"2021-06-17T20:35:20.676067Z","iopub.status.idle":"2021-06-17T20:35:20.676388Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model","metadata":{}},{"cell_type":"markdown","source":"Following starter model code is inspired from @ulrich07 <a href=\"https://www.kaggle.com/ulrich07/baseline-model-player-mean-or-median\">notebook</a> instead of just using median I am using weighted median and giving higher weight to the recent years","metadata":{}},{"cell_type":"code","source":"sample_preiction = pd.read_pickle('../input/mlb-unnested/example_sample_submission.pickle')\nexample_test_games = pd.read_pickle('../input/mlb-unnested/example_test_games.pickle')\ntrain_games = pd.read_pickle('../input/mlb-unnested/train_games.pickle')\ntrain_next_day = pd.read_pickle('../input/mlb-unnested/train_nextDayPlayerEngagement.pickle')\ntrain_next_day['year'] = pd.DatetimeIndex(train_next_day['engagementMetricsDate']).year","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-06-17T20:35:20.677110Z","iopub.status.idle":"2021-06-17T20:35:20.677419Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"year = train_next_day.year.unique()\nweight=[0.05,0.05,0.1,0.8]#Experiment here with different values","metadata":{"execution":{"iopub.status.busy":"2021-06-17T20:35:20.678130Z","iopub.status.idle":"2021-06-17T20:35:20.678446Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def weighted_median(df, val, weight):\n    df_sorted = df.sort_values(val)\n    cumsum = df_sorted[weight].cumsum()\n    cutoff = df_sorted[weight].sum() / 2.\n    return df_sorted[cumsum >= cutoff][val].iloc[0]","metadata":{"execution":{"iopub.status.busy":"2021-06-17T20:35:20.679177Z","iopub.status.idle":"2021-06-17T20:35:20.679490Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"year wise weighted targets","metadata":{}},{"cell_type":"code","source":"def preprocess(df,weight,year):\n    for i in range(len(year)):\n        df.loc[df['year']==year[i],'weight'] = weight[i]\npreprocess(train_next_day,weight,year)","metadata":{"execution":{"iopub.status.busy":"2021-06-17T20:35:20.680172Z","iopub.status.idle":"2021-06-17T20:35:20.680484Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_next_day.head()","metadata":{"execution":{"iopub.status.busy":"2021-06-17T20:35:20.681203Z","iopub.status.idle":"2021-06-17T20:35:20.681540Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"playerId = train_next_day['playerId'].unique()\nfor i in playerId:\n    df=train_next_day[train_next_day['playerId']==i]\n    wm1 = weighted_median(df,'target1','weight')\n    wm2 = weighted_median(df,'target2','weight')\n    wm3 = weighted_median(df,'target3','weight')\n    wm4 = weighted_median(df,'target4','weight')\n    train_next_day.loc[train_next_day['playerId']==i,'target1'] = wm1\n    train_next_day.loc[train_next_day['playerId']==i,'target2'] = wm2\n    train_next_day.loc[train_next_day['playerId']==i,'target3'] = wm3\n    train_next_day.loc[train_next_day['playerId']==i,'target4'] = wm4","metadata":{"execution":{"iopub.status.busy":"2021-06-17T20:35:20.682420Z","iopub.status.idle":"2021-06-17T20:35:20.682802Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_mean = train_next_day.groupby([\"playerId\"])[[\"target1\",\"target2\",\"target3\",\"target4\"]].median().reset_index()","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2021-06-17T20:35:20.683432Z","iopub.status.idle":"2021-06-17T20:35:20.683777Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_mean.head()","metadata":{"execution":{"iopub.status.busy":"2021-06-17T20:35:20.684444Z","iopub.status.idle":"2021-06-17T20:35:20.684806Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def process_pred(df):\n    df[\"playerId\"] = df[\"date_playerId\"].apply(lambda x: int( x.split(\"_\")[1] ) )\n    df.drop([\"target1\",\"target2\",\"target3\",\"target4\"], axis=1, inplace=True)\n    df = df.merge(train_mean, on=\"playerId\", how=\"left\")\n    df.drop(\"playerId\", axis=1, inplace=True)\n    df = df.fillna(0.)\n    return df","metadata":{"execution":{"iopub.status.busy":"2021-06-17T20:35:20.685489Z","iopub.status.idle":"2021-06-17T20:35:20.685832Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import mlb\nenv = mlb.make_env() # initialize the environment\niter_test = env.iter_test() # iterator which loops over each date in test set\n\nfor (test_df, sample_prediction_df) in iter_test:\n    sample_prediction_df = process_pred(sample_prediction_df)\n    env.predict(sample_prediction_df)","metadata":{"execution":{"iopub.status.busy":"2021-06-17T20:35:20.686480Z","iopub.status.idle":"2021-06-17T20:35:20.686828Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_prediction_df.head()","metadata":{"execution":{"iopub.status.busy":"2021-06-17T20:35:20.687503Z","iopub.status.idle":"2021-06-17T20:35:20.687830Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2><center>If you learned something new or forked the notebook then please don't forget to upvote<br>Thank You</center>\n</h2>","metadata":{}},{"cell_type":"markdown","source":"<h2><center>Work in Progress ... ⏳</center></h2>","metadata":{}}]}