{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Introduction and Imports 📔\n\nLet's get started with this new tabular data competition! \n\nBelow is a brief introduction to it as well as the data files we will be working with!","metadata":{}},{"cell_type":"code","source":"! pip install -q rich","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2021-06-11T04:10:15.691304Z","iopub.execute_input":"2021-06-11T04:10:15.692086Z","iopub.status.idle":"2021-06-11T04:10:25.419584Z","shell.execute_reply.started":"2021-06-11T04:10:15.691903Z","shell.execute_reply":"2021-06-11T04:10:25.418113Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport os\nimport random\n\nimport plotly.graph_objects as go\n\nfrom rich import print as _pprint","metadata":{"execution":{"iopub.status.busy":"2021-06-11T04:10:25.422454Z","iopub.execute_input":"2021-06-11T04:10:25.422868Z","iopub.status.idle":"2021-06-11T04:10:26.324103Z","shell.execute_reply.started":"2021-06-11T04:10:25.422824Z","shell.execute_reply":"2021-06-11T04:10:26.323059Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def cprint(string):\n    \"\"\"\n    Utility function for beautiful colored printing.\n    \"\"\"\n    _pprint(f\"[black]{string}[/black]\")\n\ndef show_pd_table(df, name):\n    cells_out = [df[x] for x in df.columns]\n    \n    fig = go.Figure(data=[go.Table(\n        header=dict(values=list(df.columns),\n                    fill_color='paleturquoise',\n                    align='left'),\n        cells=dict(values=cells_out,\n                   fill_color='lavender',\n                   align='left'))\n    ])\n    \n    fig.update_layout(\n    title={\n        'text': name,\n        'y':0.9,\n        'x':0.5,\n        'xanchor': 'center',\n        'yanchor': 'top'})\n\n    fig.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-06-11T04:10:26.326511Z","iopub.execute_input":"2021-06-11T04:10:26.326979Z","iopub.status.idle":"2021-06-11T04:10:26.335125Z","shell.execute_reply.started":"2021-06-11T04:10:26.326911Z","shell.execute_reply":"2021-06-11T04:10:26.333785Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## About the Competition 🏴‍☠️\n\nA player hits a walk-off home run. A pitcher throws a no-hitter. A team gets red hot going into the Postseason. We know some of the catalysts that increase baseball fan interest. Now Major League Baseball (MLB) and Google Cloud want the Kaggle community’s help to identify the many other factors which pique supporter engagement and create deeper relationships betweens players and fans.\n\nIn this competition, we will predict how fans engage with MLB players’ digital content on a daily basis for a future date range. You’ll have access to player performance data, social media data, and team factors like market size. Successful models will provide new insights into what signals most strongly correlate with and influence engagement.","metadata":{}},{"cell_type":"markdown","source":"<hr>","metadata":{}},{"cell_type":"markdown","source":"## What is our task? 🎯\n\nIn this competition, we are tasked with forecasting four different measures of engagement (`target1`-`target4`) for a subset of MLB players who are active in the 2021 season.\n\nThe data contains a set of static files that do not change with time as well as a training file of daily data which is grouped by the day.\n\nThis is a code competition that relies on a time-series module to ensure models do not peek forward in time. The time series module provides you with the test data and writes your submission file automatically. The test data arrives in a data frame identical in format to `train.csv`, except it does not contain the target values.\n\nYou are highly recommended to visit the [Evaluation](https://www.kaggle.com/c/mlb-player-digital-engagement-forecasting/overview/evaluation) Page for more info on submission.","metadata":{}},{"cell_type":"markdown","source":"<hr>","metadata":{}},{"cell_type":"markdown","source":"## How does the Data look like? 🗃\n\nSo, the data provided to us in this competition consists of 7 `.csv` files and 1 folder called `mlb/`\n\nBelow is the breakdown of the `.csv` files;\n\n* 📄 `train.csv` - This file is the training dataset.\n\n* 📄 `awards.csv` - This file is a collection of awards given out prior to the first date in the training file.\n\n* 📄 `players.csv` - This file contains high level information about all MLB players in this dataset.\n\n* 📄 `seasons.csv` - This file contains information about start and end dates of all seasons in this dataset.\n\n* 📄 `teams.csv` - This file contains high level information about all MLB teams.\n\n* 📄 `example_test.csv` - This file is an example in the form of the test set that you’ll be evaluated on..\n\n* 📄 `example_sample_submission.csv` - This file is a sample submission file in the correct format based on the example test set.","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\">\nThis is a Code comeptition and the submissions in this competition must be made using the <code>mlb</code> python module.\n</div>","metadata":{}},{"cell_type":"markdown","source":"Sample template using `mlb` module to write a submission file:\n\n```python\nimport mlb\nenv = mlb.make_env() # initialize the environment\niter_test = env.iter_test() # iterator which loops over each date in test set\n\nfor (test_df, sample_prediction_df) in iter_test:\n    sample_prediction_df['target1'] = 100 #make predictions here\n    env.predict(sample_prediction_df)\n```","metadata":{}},{"cell_type":"markdown","source":"<hr>","metadata":{}},{"cell_type":"markdown","source":"## Peeking at the Data 📈\n\nNow that you have an understanding of the task and the dataset, let's start by looking at the different data files provided and some stats on them.","metadata":{}},{"cell_type":"code","source":"awards = pd.read_csv(\"../input/mlb-player-digital-engagement-forecasting/awards.csv\")\nplayers = pd.read_csv(\"../input/mlb-player-digital-engagement-forecasting/players.csv\")\nseasons = pd.read_csv(\"../input/mlb-player-digital-engagement-forecasting/seasons.csv\")\nteams = pd.read_csv(\"../input/mlb-player-digital-engagement-forecasting/teams.csv\")\ntrain = pd.read_csv(\"../input/mlb-player-digital-engagement-forecasting/train.csv\")","metadata":{"execution":{"iopub.status.busy":"2021-06-10T15:47:09.373685Z","iopub.execute_input":"2021-06-10T15:47:09.373957Z","iopub.status.idle":"2021-06-10T15:48:20.044139Z","shell.execute_reply.started":"2021-06-10T15:47:09.37393Z","shell.execute_reply":"2021-06-10T15:48:20.043342Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.head()","metadata":{"execution":{"iopub.status.busy":"2021-06-10T15:48:20.045195Z","iopub.execute_input":"2021-06-10T15:48:20.045542Z","iopub.status.idle":"2021-06-10T15:48:20.08971Z","shell.execute_reply.started":"2021-06-10T15:48:20.045517Z","shell.execute_reply":"2021-06-10T15:48:20.088885Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"show_pd_table(awards.head(20), name='Awards DataFrame [First-20 Values]')","metadata":{"execution":{"iopub.status.busy":"2021-06-10T15:48:20.090854Z","iopub.execute_input":"2021-06-10T15:48:20.09112Z","iopub.status.idle":"2021-06-10T15:48:20.248492Z","shell.execute_reply.started":"2021-06-10T15:48:20.091093Z","shell.execute_reply":"2021-06-10T15:48:20.247438Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"show_pd_table(players.head(20), name='Players DataFrame [First-20 Values]')","metadata":{"execution":{"iopub.status.busy":"2021-06-10T15:48:20.249752Z","iopub.execute_input":"2021-06-10T15:48:20.250035Z","iopub.status.idle":"2021-06-10T15:48:20.272721Z","shell.execute_reply.started":"2021-06-10T15:48:20.250008Z","shell.execute_reply":"2021-06-10T15:48:20.271621Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"show_pd_table(seasons, name='Seasons DataFrame [First-20 Values]')","metadata":{"execution":{"iopub.status.busy":"2021-06-10T15:48:20.274822Z","iopub.execute_input":"2021-06-10T15:48:20.275134Z","iopub.status.idle":"2021-06-10T15:48:20.291083Z","shell.execute_reply.started":"2021-06-10T15:48:20.275102Z","shell.execute_reply":"2021-06-10T15:48:20.290397Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"show_pd_table(teams.head(20), name='Teams DataFrame [First-20 Values]')","metadata":{"execution":{"iopub.status.busy":"2021-06-10T15:48:20.292188Z","iopub.execute_input":"2021-06-10T15:48:20.292479Z","iopub.status.idle":"2021-06-10T15:48:20.312286Z","shell.execute_reply.started":"2021-06-10T15:48:20.292448Z","shell.execute_reply":"2021-06-10T15:48:20.31122Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"players.describe()","metadata":{"execution":{"iopub.status.busy":"2021-06-10T15:54:54.078727Z","iopub.execute_input":"2021-06-10T15:54:54.079055Z","iopub.status.idle":"2021-06-10T15:54:54.098833Z","shell.execute_reply.started":"2021-06-10T15:54:54.07903Z","shell.execute_reply":"2021-06-10T15:54:54.097769Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Player Height and Weight Distribution","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(nrows=1, ncols=2, figsize=(12, 7))\nfig.suptitle(f\"Height and Weight Distribution\")\n\nsns.histplot(players['heightInches'], stat='density', color='magenta', legend=True, ax=ax[0])\nsns.histplot(players['weight'], stat='density', color='blue', legend=True, ax=ax[1])\n\nax[0].set_xlabel(\"Height\")\nax[1].set_xlabel(\"Weight\")\nax[0].set_ylabel(\"Density\")\nax[1].set_ylabel(\"Density\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2021-06-10T16:09:03.895366Z","iopub.execute_input":"2021-06-10T16:09:03.895711Z","iopub.status.idle":"2021-06-10T16:09:04.390623Z","shell.execute_reply.started":"2021-06-10T16:09:03.895683Z","shell.execute_reply":"2021-06-10T16:09:04.388966Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"_pprint(\"[bold green]Under Work![/bold green]\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-06-11T04:10:38.253156Z","iopub.execute_input":"2021-06-11T04:10:38.253557Z","iopub.status.idle":"2021-06-11T04:10:38.260448Z","shell.execute_reply.started":"2021-06-11T04:10:38.253523Z","shell.execute_reply":"2021-06-11T04:10:38.259388Z"},"trusted":true},"execution_count":null,"outputs":[]}]}