{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Getting Started with MLB Player Digital Engagement Forecasting #","metadata":{"execution":{"iopub.execute_input":"2021-06-08T23:38:55.536369Z","iopub.status.busy":"2021-06-08T23:38:55.536027Z","iopub.status.idle":"2021-06-08T23:38:55.540097Z","shell.execute_reply":"2021-06-08T23:38:55.53902Z","shell.execute_reply.started":"2021-06-08T23:38:55.536341Z"},"papermill":{"duration":0.023885,"end_time":"2021-06-18T03:20:01.694054","exception":false,"start_time":"2021-06-18T03:20:01.670169","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"Welcome to the Getting Started notebook for the *MLB Digital Engagement Forecasting* competition! This notebook is a complete start-to-finish guide to entering this competition. We will:\n- load and join the data,\n- explore time series properties,\n- create a feature set,\n- train a neural network, and\n- make a submission.\n\nAs this is a *code competition*, be aware that your submission will be a notebook to make predictions during the test period. Read more about [code competitions](https://www.kaggle.com/docs/competitions#notebooks-only-competitions).\n\nIn the complementary notebook, [Vertex AI with MLB Player Digital Engagement](https://www.kaggle.com/ryanholbrook/vertex-ai-with-mlb-player-digital-engagement), we'll also demonstrate some of the capabilities of **Vertex AI**, Google's new unified AI platform:\n- using **Vertex AI Notebooks** on Google Cloud Platform\n- exploring **Explainable AI** on Vertex AI to refine your features\n- hyperparameter tuning with **Vertex Vizier**","metadata":{"papermill":{"duration":0.022268,"end_time":"2021-06-18T03:20:01.739158","exception":false,"start_time":"2021-06-18T03:20:01.716890","status":"completed"},"tags":[]}},{"cell_type":"code","source":"import gc\nimport sys\nimport warnings\nfrom joblib import Parallel, delayed\nfrom pathlib import Path\n\nimport ipywidgets as widgets\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport pandas as pd\nimport seaborn as sns\nfrom sklearn.model_selection import train_test_split\nfrom statsmodels.tsa.deterministic import (CalendarFourier,\n                                           CalendarSeasonality,\n                                           CalendarTimeTrend,\n                                           DeterministicProcess)\n\nimport tensorflow as tf\nfrom tensorflow import keras\nfrom tensorflow.keras import layers\nfrom keras.layers.experimental.preprocessing import StringLookup\n\nwarnings.simplefilter(\"ignore\")\n\n# Set Matplotlib defaults\nplt.style.use(\"seaborn-whitegrid\")\nplt.rc(\"figure\", autolayout=True, figsize=(11, 5))\nplt.rc(\n    \"axes\",\n    labelweight=\"bold\",\n    labelsize=\"large\",\n    titleweight=\"bold\",\n    titlesize=14,\n    titlepad=10,\n)\nplot_params = dict(\n    color=\"0.75\",\n    style=\".-\",\n    markeredgecolor=\"0.25\",\n    markerfacecolor=\"0.25\",\n    legend=False,\n)","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-input":true,"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","papermill":{"duration":7.400126,"end_time":"2021-06-18T03:20:09.162691","exception":false,"start_time":"2021-06-18T03:20:01.762565","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2021-06-18T14:19:39.340573Z","iopub.execute_input":"2021-06-18T14:19:39.340971Z","iopub.status.idle":"2021-06-18T14:19:48.062312Z","shell.execute_reply.started":"2021-06-18T14:19:39.340887Z","shell.execute_reply":"2021-06-18T14:19:48.061062Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Create Training Data #","metadata":{"papermill":{"duration":0.021863,"end_time":"2021-06-18T03:20:09.206804","exception":false,"start_time":"2021-06-18T03:20:09.184941","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"There is a lot more data available than what we'll use in this notebook. See the [data documentation](https://www.kaggle.com/c/mlb-player-digital-engagement-forecasting/data) on the competition page for a complete description the *MLB Player Digital Engagement Forecasting* competition data. And be sure to check out Google data scientist Alok Pattani's in depth data exploration: [MLB Player Digital Engagement Data Exploration](https://www.kaggle.com/alokpattani/mlb-player-digital-engagement-data-exploration).","metadata":{"papermill":{"duration":0.021924,"end_time":"2021-06-18T03:20:09.256426","exception":false,"start_time":"2021-06-18T03:20:09.234502","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"### Read and extract dataframes","metadata":{"papermill":{"duration":0.021867,"end_time":"2021-06-18T03:20:09.300402","exception":false,"start_time":"2021-06-18T03:20:09.278535","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Helper function to unpack json found in daily data\ndef unpack_json(json_str):\n    return pd.DataFrame() if pd.isna(json_str) else pd.read_json(json_str)\n\n\ndef unpack_data(data, dfs=None, n_jobs=-1):\n    if dfs is not None:\n        data = data.loc[:, dfs]\n    unnested_dfs = {}\n    for name, column in data.iteritems():\n        daily_dfs = Parallel(n_jobs=n_jobs)(\n            delayed(unpack_json)(item) for date, item in column.iteritems())\n        df = pd.concat(daily_dfs)\n        unnested_dfs[name] = df\n    return unnested_dfs","metadata":{"_kg_hide-input":true,"papermill":{"duration":0.032719,"end_time":"2021-06-18T03:20:09.355306","exception":false,"start_time":"2021-06-18T03:20:09.322587","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2021-06-18T14:19:48.063829Z","iopub.execute_input":"2021-06-18T14:19:48.064121Z","iopub.status.idle":"2021-06-18T14:19:48.071913Z","shell.execute_reply.started":"2021-06-18T14:19:48.064089Z","shell.execute_reply":"2021-06-18T14:19:48.070140Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are a number of supplementary files in addition to the training data.","metadata":{"papermill":{"duration":0.021999,"end_time":"2021-06-18T03:20:09.399627","exception":false,"start_time":"2021-06-18T03:20:09.377628","status":"completed"},"tags":[]}},{"cell_type":"code","source":"data_dir = Path('../input/mlb-player-digital-engagement-forecasting/')\n\ndf_names = ['seasons', 'teams', 'players', 'awards']\n\nfor name in df_names:\n    globals()[name] = pd.read_csv(data_dir / f\"{name}.csv\")\n\nkaggle_data_tabs = widgets.Tab()\n# Add Output widgets for each pandas DF as tabs' children\nkaggle_data_tabs.children = list([widgets.Output() for df_name in df_names])\n\nfor index in range(0, len(df_names)):\n    # Rename tab bar titles to df names\n    kaggle_data_tabs.set_title(index, df_names[index])\n    \n    # Display corresponding table output for this tab name\n    with kaggle_data_tabs.children[index]:\n        display(eval(df_names[index]))\n\ndisplay(kaggle_data_tabs)","metadata":{"_kg_hide-input":true,"papermill":{"duration":0.243274,"end_time":"2021-06-18T03:20:09.665119","exception":false,"start_time":"2021-06-18T03:20:09.421845","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2021-06-18T14:19:48.074308Z","iopub.execute_input":"2021-06-18T14:19:48.074793Z","iopub.status.idle":"2021-06-18T14:19:48.301021Z","shell.execute_reply.started":"2021-06-18T14:19:48.074762Z","shell.execute_reply":"2021-06-18T14:19:48.299616Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The training data is a time-indexed collection of nested JSON fields containing information about each player. The targets are contained in the `nextDayPlayerEngagement` column, while the remaining columns contain data you could use to construct features. For this getting started notebook, we'll only use features from the `playerBoxScores` dataframe.","metadata":{"papermill":{"duration":0.022534,"end_time":"2021-06-18T03:20:09.713647","exception":false,"start_time":"2021-06-18T03:20:09.691113","status":"completed"},"tags":[]}},{"cell_type":"code","source":"%%time\n# Define dataframes to load from training set\ndfs = [\n    'nextDayPlayerEngagement',  # targets\n    'playerBoxScores',  # features\n    # Other dataframes available for features:\n    # 'games',\n    # 'rosters',\n    # 'teamBoxScores',\n    # 'transactions',\n    # 'standings',\n    # 'awards',\n    # 'events',\n    # 'playerTwitterFollowers',\n    # 'teamTwitterFollowers',\n]\n\n# Read training data\ntraining = pd.read_csv(\n    data_dir / 'train.csv',\n    usecols=['date'] + dfs,\n)\n\n# Convert training data date field to datetime type\ntraining['date'] = pd.to_datetime(training['date'], format=\"%Y%m%d\")\ntraining = training.set_index('date').to_period('D')\nprint(training.info())","metadata":{"_kg_hide-input":false,"papermill":{"duration":64.466592,"end_time":"2021-06-18T03:21:14.203064","exception":false,"start_time":"2021-06-18T03:20:09.736472","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2021-06-18T14:19:48.303457Z","iopub.execute_input":"2021-06-18T14:19:48.303796Z","iopub.status.idle":"2021-06-18T14:20:57.266823Z","shell.execute_reply.started":"2021-06-18T14:19:48.303772Z","shell.execute_reply":"2021-06-18T14:20:57.265271Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%time\n# Unpack nested dataframes and store in dictionary `training_dfs`\ntraining_dfs = unpack_data(training, dfs=dfs)\nprint('\\n', training_dfs.keys())","metadata":{"papermill":{"duration":26.236092,"end_time":"2021-06-18T03:21:40.462615","exception":false,"start_time":"2021-06-18T03:21:14.226523","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2021-06-18T14:20:57.268823Z","iopub.execute_input":"2021-06-18T14:20:57.269353Z","iopub.status.idle":"2021-06-18T14:21:16.906533Z","shell.execute_reply.started":"2021-06-18T14:20:57.269300Z","shell.execute_reply":"2021-06-18T14:21:16.905191Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Construct training and validation splits","metadata":{"papermill":{"duration":0.023241,"end_time":"2021-06-18T03:21:40.509814","exception":false,"start_time":"2021-06-18T03:21:40.486573","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"Defined in the next cell are a number of functions that will process our data into training and validation splits.","metadata":{}},{"cell_type":"code","source":"# Players in the test set. We'll filter our data for only this set of players\npids_test = players.playerId.loc[\n    players.playerForTestSetAndFuturePreds.fillna(False)\n].astype(str)\n\n# Name of target columns\ntargets = [\"target1\", \"target2\", \"target3\", \"target4\"]\n\n\ndef make_playerBoxScores(dfs: dict, features):\n    X = dfs['playerBoxScores'].copy()\n    X = X[['gameDate', 'playerId'] + features]\n    # Set dtypes\n    X = X.astype({name: np.float32 for name in features})\n    X = X.astype({'playerId': str})\n    # Create date index\n    X = X.rename(columns={'gameDate': 'date'})\n    X['date'] = pd.PeriodIndex(X.date, freq='D')\n    # Aggregate multiple games per day by summing\n    X = X.groupby(['date', 'playerId'], as_index=False).sum()\n    return X\n\n\ndef make_targets(training_dfs: dict):\n    Y = training_dfs['nextDayPlayerEngagement'].copy()\n    # Set dtypes\n    Y = Y.astype({name: np.float32 for name in targets})\n    Y = Y.astype({'playerId': str})\n    # Match target dates to feature dates and create date index\n    Y = Y.rename(columns={'engagementMetricsDate': 'date'})\n    Y['date'] = pd.to_datetime(Y['date'])\n    Y = Y.set_index('date').to_period('D')\n    Y.index = Y.index - 1\n    return Y.reset_index()\n\n\ndef join_datasets(dfs):\n    dfs = [x.pivot(index='date', columns='playerId') for x in dfs]\n    df = pd.concat(dfs, axis=1).stack().reset_index('playerId')\n    return df\n\n\ndef make_training_data(training_dfs: dict,\n                       features,\n                       targets,\n                       fourier=4,\n                       test_size=30):\n    # Process dataframes\n    X = make_playerBoxScores(training_dfs, features)\n    Y = make_targets(training_dfs)\n    # Merge for processing\n    df = join_datasets([X, Y])\n    # Filter for players in test set\n    df = df.loc[df.playerId.isin(pids_test), :]\n    # Convert from long to wide format\n    df = df.pivot(columns=\"playerId\")\n    # Restore features and targets\n    X = df.loc(axis=1)[features, :]\n    Y = df.loc(axis=1)[targets, :]\n    # Fill missing values in features\n    X.fillna(-1, inplace=True)\n    # Create temporal features\n    fourier_terms = CalendarFourier(freq='A', order=fourier)\n    deterministic = DeterministicProcess(\n        index=X.index,\n        order=0,\n        seasonal=False,  # set to True for weekly seasonality\n        additional_terms=[fourier_terms],\n    )\n    X = pd.concat([X, deterministic.in_sample()], axis=1)\n    # Create train / validation splits\n    X_train, X_valid, y_train, y_valid = train_test_split(\n        X,\n        Y,\n        test_size=test_size,\n        shuffle=False,\n    )\n    return X_train, X_valid, y_train, y_valid, deterministic","metadata":{"_kg_hide-input":true,"papermill":{"duration":0.047912,"end_time":"2021-06-18T03:21:40.581553","exception":false,"start_time":"2021-06-18T03:21:40.533641","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2021-06-18T14:21:16.908183Z","iopub.execute_input":"2021-06-18T14:21:16.908463Z","iopub.status.idle":"2021-06-18T14:21:16.925779Z","shell.execute_reply.started":"2021-06-18T14:21:16.908438Z","shell.execute_reply":"2021-06-18T14:21:16.924801Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In this cell we'll define the features we want to use from `playerBoxScores`. This frame contains player statistics for every game in the training period. We've chosen just a few from the many available. The `make_training_data` function also adjoins some features modeling annual seasonality, motivated by our data exploration in the next section.","metadata":{}},{"cell_type":"code","source":"%%time\n# Columns to select from playerBoxScores, all numeric\nfeatures = [\n    \"hits\",\n    \"strikeOuts\",\n    \"homeRuns\",\n    \"runsScored\",\n    \"stolenBases\",\n    \"strikeOutsPitching\",\n    \"inningsPitched\",\n    \"strikes\",\n    \"flyOuts\",\n    \"groundOuts\",\n    \"errors\",\n]\n\n# Number of days to use for the validation set\ntest_size = 30\n\nX_train, X_valid, y_train, y_valid, deterministic = make_training_data(\n    training_dfs, \n    features=features, \n    targets=targets,\n    fourier=4,  # number of Fourier pairs describing annual seasonality\n    test_size=test_size,\n)","metadata":{"papermill":{"duration":28.066727,"end_time":"2021-06-18T03:22:08.672119","exception":false,"start_time":"2021-06-18T03:21:40.605392","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2021-06-18T14:21:16.927126Z","iopub.execute_input":"2021-06-18T14:21:16.927472Z","iopub.status.idle":"2021-06-18T14:21:38.018139Z","shell.execute_reply.started":"2021-06-18T14:21:16.927442Z","shell.execute_reply":"2021-06-18T14:21:38.016510Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Exploration #","metadata":{"papermill":{"duration":0.023802,"end_time":"2021-06-18T03:22:08.720276","exception":false,"start_time":"2021-06-18T03:22:08.696474","status":"completed"},"tags":[]}},{"cell_type":"code","source":"def seasonal_plot(X, y, period, freq, ax=None):\n    if ax is None:\n        _, ax = plt.subplots()\n    palette = sns.color_palette(\n        \"husl\",\n        n_colors=X[period].nunique(),\n    )\n    ax = sns.lineplot(\n        x=freq,\n        y=y,\n        hue=period,\n        data=X,\n        ci=False,\n        ax=ax,\n        palette=palette,\n        legend=False,\n    )\n    ax.set_title(f\"Seasonal Plot ({period}/{freq})\")\n    for line, name in zip(ax.lines, X[period].unique()):\n        y_ = line.get_ydata()[-1]\n        ax.annotate(\n            name,\n            xy=(1, y_),\n            xytext=(6, 0),\n            color=line.get_color(),\n            xycoords=ax.get_yaxis_transform(),\n            textcoords=\"offset points\",\n            size=14,\n            va=\"center\",\n        )\n    return ax\n\n\ndef plot_periodogram(ts, detrend='linear', ax=None):\n    from scipy.signal import periodogram\n    fs = pd.Timedelta(\"1Y\") / pd.Timedelta(\"1D\")\n    freqencies, spectrum = periodogram(\n        ts,\n        fs=fs,\n        detrend=detrend,\n        window=\"boxcar\",\n        scaling='spectrum',\n    )\n    if ax is None:\n        _, ax = plt.subplots()\n    ax.step(freqencies, spectrum, color=\"purple\")\n    ax.set_xscale(\"log\")\n    ax.set_xticks([1, 2, 4, 6, 12, 26, 52, 104])\n    ax.set_xticklabels(\n        [\n            \"Annual\",\n            \"Semiannual\",\n            \"Quarterly\",\n            \"Bimonthly\",\n            \"Monthly\",\n            \"Biweekly\",\n            \"Weekly\",\n            \"Semiweekly\",\n        ],\n        rotation=30,\n    )\n    ax.ticklabel_format(axis=\"y\", style=\"sci\", scilimits=(0, 0))\n    ax.set_ylabel(\"Density\")\n    ax.set_title(\"Periodogram\")\n    return ax","metadata":{"_kg_hide-input":true,"papermill":{"duration":0.042772,"end_time":"2021-06-18T03:22:08.837530","exception":false,"start_time":"2021-06-18T03:22:08.794758","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2021-06-18T14:21:38.021412Z","iopub.execute_input":"2021-06-18T14:21:38.021763Z","iopub.status.idle":"2021-06-18T14:21:38.035205Z","shell.execute_reply.started":"2021-06-18T14:21:38.021736Z","shell.execute_reply":"2021-06-18T14:21:38.033311Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"A *seasonal plot* can reveal seasonal effects in a time series. Let's look at seasonal plots for players in the top 10% of average engagement, one for yearly seasonality and one for weekly seasonality.","metadata":{"papermill":{"duration":0.023666,"end_time":"2021-06-18T03:22:08.885309","exception":false,"start_time":"2021-06-18T03:22:08.861643","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Select players in top decile of engagement\ndeciles = pd.qcut(y_train.mean().mean(level=1), q=10)\npids_top_decile = deciles.index[deciles == deciles.max()]\ny_top_decile = y_train.loc(axis=1)[:, pids_top_decile]\n\n# Create average engagement series\ny_top_decile_avg = (y_top_decile / y_top_decile.max(axis=0)).mean(axis=1)\ny_top_decile_avg.name = \"target\"\n\n# Yearly plot\nS = y_top_decile_avg.to_frame()\nS[\"month\"] = S.index.month  # the frequency\nS[\"year\"] = S.index.year  # the period\n_ = seasonal_plot(S, y=\"target\", period=\"year\", freq=\"month\")\n\n# Weekly plot\nS = y_top_decile_avg.to_frame()\nS[\"day\"] = S.index.dayofweek  # the frequency\nS[\"week\"] = S.index.week  # the period\n_ = seasonal_plot(S, y=\"target\", period=\"week\", freq=\"day\")","metadata":{"_kg_hide-input":true,"papermill":{"duration":13.861242,"end_time":"2021-06-18T03:22:22.770592","exception":false,"start_time":"2021-06-18T03:22:08.909350","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2021-06-18T14:21:38.037340Z","iopub.execute_input":"2021-06-18T14:21:38.037680Z","iopub.status.idle":"2021-06-18T14:21:46.699900Z","shell.execute_reply.started":"2021-06-18T14:21:38.037654Z","shell.execute_reply":"2021-06-18T14:21:46.698392Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From these plots, digital player engagement appears to have a strong annual component, but little or no weekly component.\n\nWe can verify this with the *periodogram*. The periodogram illustrates the strength of the frequencies within a signal -- specifically, the variance of the sine / cosine Fourier component oscillating at that frequency.","metadata":{"papermill":{"duration":0.033913,"end_time":"2021-06-18T03:22:22.841311","exception":false,"start_time":"2021-06-18T03:22:22.807398","status":"completed"},"tags":[]}},{"cell_type":"code","source":"_ = plot_periodogram(y_top_decile_avg)","metadata":{"_kg_hide-input":true,"papermill":{"duration":0.564637,"end_time":"2021-06-18T03:22:23.446789","exception":false,"start_time":"2021-06-18T03:22:22.882152","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2021-06-18T14:21:46.702007Z","iopub.execute_input":"2021-06-18T14:21:46.702606Z","iopub.status.idle":"2021-06-18T14:21:47.107854Z","shell.execute_reply.started":"2021-06-18T14:21:46.702566Z","shell.execute_reply":"2021-06-18T14:21:47.106364Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Both of these visualizations indicate yearly (or annual) seasonality in the engagement time series, which motivated our decision to use yearly Fourier features when constructing our feature set.","metadata":{"papermill":{"duration":0.034724,"end_time":"2021-06-18T03:22:23.522898","exception":false,"start_time":"2021-06-18T03:22:23.488174","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"# Model #","metadata":{"papermill":{"duration":0.034856,"end_time":"2021-06-18T03:22:23.592867","exception":false,"start_time":"2021-06-18T03:22:23.558011","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"For this getting started notebook, we'll just use a simple feedforward network.","metadata":{"papermill":{"duration":0.034794,"end_time":"2021-06-18T03:22:23.662639","exception":false,"start_time":"2021-06-18T03:22:23.627845","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Hyperparameters\nHIDDEN = 1024\nACTIVATION = 'relu'  # could try elu, gelu, swish\nDROPOUT_RATE = 0.5\nLEARNING_RATE = 1e-2\nBATCH_SIZE = 32\n\nOUTPUTS = y_train.shape[-1]\nmodel = keras.Sequential([\n    layers.Dense(HIDDEN, activation=ACTIVATION),\n    layers.BatchNormalization(),\n    layers.Dropout(DROPOUT_RATE),\n    layers.Dense(HIDDEN, activation=ACTIVATION),\n    layers.BatchNormalization(),\n    layers.Dropout(DROPOUT_RATE),\n    layers.Dense(HIDDEN, activation=ACTIVATION),\n    layers.BatchNormalization(),\n    layers.Dropout(DROPOUT_RATE),\n    layers.Dense(OUTPUTS),\n])","metadata":{"papermill":{"duration":0.095478,"end_time":"2021-06-18T03:22:23.794033","exception":false,"start_time":"2021-06-18T03:22:23.698555","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2021-06-18T14:21:47.109478Z","iopub.execute_input":"2021-06-18T14:21:47.109764Z","iopub.status.idle":"2021-06-18T14:21:47.214099Z","shell.execute_reply.started":"2021-06-18T14:21:47.109739Z","shell.execute_reply":"2021-06-18T14:21:47.212589Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\noptimizer = keras.optimizers.Adam(learning_rate=LEARNING_RATE)\nmodel.compile(optimizer=optimizer, loss='mae', metrics=['mae'])\n\nearly_stopping = keras.callbacks.EarlyStopping(patience=3)\n\nhistory = model.fit(\n    X_train, y_train,\n    validation_data=(X_valid, y_valid),\n    batch_size=BATCH_SIZE,\n    epochs=100,\n    callbacks=[early_stopping],\n)","metadata":{"papermill":{"duration":58.423254,"end_time":"2021-06-18T03:23:22.253220","exception":false,"start_time":"2021-06-18T03:22:23.829966","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2021-06-18T14:21:47.215895Z","iopub.execute_input":"2021-06-18T14:21:47.216151Z","iopub.status.idle":"2021-06-18T14:22:32.675472Z","shell.execute_reply.started":"2021-06-18T14:21:47.216125Z","shell.execute_reply":"2021-06-18T14:22:32.674366Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Create Submission #","metadata":{"papermill":{"duration":0.188102,"end_time":"2021-06-18T03:23:22.630419","exception":false,"start_time":"2021-06-18T03:23:22.442317","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"The next cell illustrates how you to create a submission for this competition. As this is a code competition that relies on a time series module, submissions must follow the requirements described on the [Evaluation Page](https://www.kaggle.com/c/mlb-player-digital-engagement-forecasting/overview/evaluation) and [Data Page](https://www.kaggle.com/c/mlb-player-digital-engagement-forecasting/data).","metadata":{"papermill":{"duration":0.188663,"end_time":"2021-06-18T03:23:23.007715","exception":false,"start_time":"2021-06-18T03:23:22.819052","status":"completed"},"tags":[]}},{"cell_type":"code","source":"def make_test_data(test_dfs: dict, features, deterministic):\n    X = make_playerBoxScores(test_dfs, features)\n    X = X.merge(pids_test, how='right')\n    X['date'] = X.date.fillna(method='ffill').fillna(method='bfill')\n    X.fillna(-1, inplace=True)\n    # Convert from long to wide format\n    X = X.pivot(index='date', columns=\"playerId\")\n    # Create temporal features\n    X = pd.concat([\n        X,\n        deterministic.out_of_sample(steps=1, forecast_index=X.index),\n    ],\n                  axis=1)\n    return X\n\n\ndef make_predictions(model, X, columns, targets):\n    y_pred = model.predict(X)\n    y_pred = pd.DataFrame(y_pred, columns=columns, index=X.index).stack()\n    y_pred[targets] = y_pred[targets].clip(0, 100)\n    y_pred['date_playerId'] = [\n        (date + 1).strftime('%Y%m%d') + '_' + str(playerId)\n        for date, playerId in y_pred.index\n    ]\n    y_pred.reset_index('playerId', drop=True, inplace=True)\n    y_pred = y_pred[['date_playerId'] + targets]  # reorder\n    y_pred.index = pd.Int64Index(\n        [int(date.strftime('%Y%m%d')) for date in y_pred.index], name='date')\n    return y_pred","metadata":{"_kg_hide-input":true,"papermill":{"duration":0.20493,"end_time":"2021-06-18T03:23:23.401975","exception":false,"start_time":"2021-06-18T03:23:23.197045","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2021-06-18T14:22:32.677310Z","iopub.execute_input":"2021-06-18T14:22:32.677667Z","iopub.status.idle":"2021-06-18T14:22:32.688204Z","shell.execute_reply.started":"2021-06-18T14:22:32.677633Z","shell.execute_reply":"2021-06-18T14:22:32.687370Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nimport mlb\n\nenv = mlb.make_env()\niter_test = env.iter_test()\n\nfor (test_df, sample_prediction_df) in iter_test:\n    # Unpack features from test_df\n    test_dfs = unpack_data(test_df, dfs=['playerBoxScores'])\n    X = make_test_data(test_dfs, features, deterministic)\n\n    # Create predictions\n    y_pred = make_predictions(\n        model,\n        X,\n        columns=y_train.columns,\n        targets=targets,\n    )\n    submission = (\n        sample_prediction_df\n        [['date_playerId']]\n        .reset_index()  #  preserve index 'date'\n        .merge(y_pred, how='left', on='date_playerId')\n        .set_index('date')  #  restore index 'date'\n    )\n\n    # Submit predictions\n    env.predict(submission)  # constructs submissions.csv","metadata":{"papermill":{"duration":3.576003,"end_time":"2021-06-18T03:23:27.166726","exception":false,"start_time":"2021-06-18T03:23:23.590723","status":"completed"},"scrolled":true,"tags":[],"execution":{"iopub.status.busy":"2021-06-18T14:22:32.689539Z","iopub.execute_input":"2021-06-18T14:22:32.689861Z","iopub.status.idle":"2021-06-18T14:22:35.503202Z","shell.execute_reply.started":"2021-06-18T14:22:32.689826Z","shell.execute_reply":"2021-06-18T14:22:35.501412Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"To complete a submission for this competition, we'll need to commit this notebook and submit the resulting submission file it creates,`submissions.csv`. From the notebook editor:\n- make sure *Internet* is turned off in **Settings**,\n- click the **Save Version** button to the upper right,\n- make sure *Save and Run All (Commit)* is selected, and\n- click **Save**.\n\nOnce the commit completes (should only be three or four minutes):\n- select from the menubar **File -> Version History**,\n- select from the ellipsis menu of the latest version **... -> Submit to Competition**, and\n- click **Submit**\n\nKaggle will rerun the notebook on the public test set and display the score on the public leaderboard. [**My Submissions**](https://www.kaggle.com/c/mlb-player-digital-engagement-forecasting/submissions) will show the status of the submission: currently running, succeeded, or failed.\n\nThe submission requirements for code competitions can be somewhat exacting. If you're having trouble, checkout our [**Code Competitions - Errors & Debugging Tips**](https://www.kaggle.com/docs/competitions#notebooks-only-FAQ).","metadata":{"papermill":{"duration":0.219759,"end_time":"2021-06-18T03:23:27.982104","exception":false,"start_time":"2021-06-18T03:23:27.762345","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"# Explainable AI and Hyperparameter Tuning with Vertex #\n\n[This complementary notebook](https://www.kaggle.com/ryanholbrook/vertex-ai-with-mlb-player-digital-engagement) demonstrates how to run this notebook in Vertex AI Notebooks. You'll learn how to use **Explainable AI (XAI)** to refine your features and how to tune your model with **Vertex Vizier**.","metadata":{"papermill":{"duration":0.192104,"end_time":"2021-06-18T03:23:28.387072","exception":false,"start_time":"2021-06-18T03:23:28.194968","status":"completed"},"tags":[]}}]}