{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.12.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":27783,"databundleVersionId":2555459,"sourceType":"competition"},{"sourceId":14205403,"sourceType":"datasetVersion","datasetId":9060199}],"dockerImageVersionId":31234,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"import os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\nfrom plotly.offline import init_notebook_mode, iplot\ninit_notebook_mode(connected=True)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-12-18T06:11:36.721325Z","iopub.execute_input":"2025-12-18T06:11:36.721843Z","iopub.status.idle":"2025-12-18T06:11:36.755571Z","shell.execute_reply.started":"2025-12-18T06:11:36.721807Z","shell.execute_reply":"2025-12-18T06:11:36.754294Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from prophet import Prophet\nimport pandas as pd\nimport numpy as np\nimport json\nimport matplotlib.pyplot as plt\nimport seaborn as sns","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-18T04:22:05.228941Z","iopub.execute_input":"2025-12-18T04:22:05.230059Z","iopub.status.idle":"2025-12-18T04:22:07.401786Z","shell.execute_reply.started":"2025-12-18T04:22:05.229980Z","shell.execute_reply":"2025-12-18T04:22:07.400789Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Objective\n\nThe goal of this project is to forecast daily MLB digital engagement at the league level using the datasets from Kaggle’s MLB Player Digital Engagement Forecasting competition. Since the engagement targets are reported for the next day, I align the timeline so ds represents the day the engagement is measured. From there, I build an end-to-end workflow including data preprocessing, EDA to understand seasonality and weekly patterns, training a Prophet forecasting model, and model evaluation with baseline comparisons. Overall, this forecast can help MLB teams plan ahead by flagging upcoming high-engagement periods so teams can schedule content, campaigns, and staffing more strategically.","metadata":{}},{"cell_type":"code","source":"df = pd.read_csv(\"/kaggle/input/mlb-player-digital-engagement-forecasting/train.csv\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-18T04:24:41.542649Z","iopub.execute_input":"2025-12-18T04:24:41.543266Z","iopub.status.idle":"2025-12-18T04:26:33.455674Z","shell.execute_reply.started":"2025-12-18T04:24:41.543226Z","shell.execute_reply":"2025-12-18T04:26:33.454664Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-18T04:26:33.457100Z","iopub.execute_input":"2025-12-18T04:26:33.457397Z","iopub.status.idle":"2025-12-18T04:26:33.500443Z","shell.execute_reply.started":"2025-12-18T04:26:33.457366Z","shell.execute_reply":"2025-12-18T04:26:33.499395Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Data Preprocessing\n\nSince this dataset provides four engagement targets for each player, I first take the daily average across all players for each target, which gives me y1, y2, y3, and y4. Then I combine these four values into a single metric y, which represents the overall league-wide daily engagement (averaged across targets and across players) for each day.","metadata":{}},{"cell_type":"code","source":"df[\"ds\"] = pd.to_datetime(df[\"date\"].astype(str), format = \"%Y%m%d\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-18T04:27:09.587926Z","iopub.execute_input":"2025-12-18T04:27:09.588320Z","iopub.status.idle":"2025-12-18T04:27:09.599629Z","shell.execute_reply.started":"2025-12-18T04:27:09.588288Z","shell.execute_reply":"2025-12-18T04:27:09.598647Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def agg_next_day_targets(cell, how=\"mean\"):\n    \"\"\"\n    cell: JSON string for nextDayPlayerEngagement (list of dicts)\n    returns: (y1,y2,y3,y4,n_players)\n    \"\"\"\n    if pd.isna(cell) or cell == \"\":\n        return pd.Series([np.nan]*5)\n\n    arr = json.loads(cell)  # list of dicts\n    if not arr:\n        return pd.Series([np.nan]*5)\n\n    t1 = np.array([d.get(\"target1\", np.nan) for d in arr], dtype=float)\n    t2 = np.array([d.get(\"target2\", np.nan) for d in arr], dtype=float)\n    t3 = np.array([d.get(\"target3\", np.nan) for d in arr], dtype=float)\n    t4 = np.array([d.get(\"target4\", np.nan) for d in arr], dtype=float)\n\n    if how == \"sum\":\n        f = np.nansum\n    else:\n        f = np.nanmean\n\n    return pd.Series([f(t1), f(t2), f(t3), f(t4)])\n\ndf[[\"y1\",\"y2\",\"y3\",\"y4\"]] = df[\"nextDayPlayerEngagement\"].apply(agg_next_day_targets)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-18T04:27:17.505535Z","iopub.execute_input":"2025-12-18T04:27:17.505981Z","iopub.status.idle":"2025-12-18T04:27:27.422976Z","shell.execute_reply.started":"2025-12-18T04:27:17.505942Z","shell.execute_reply":"2025-12-18T04:27:27.421997Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Building Regressors\n\nSince Prophet only requires ds and y, it can already model trend and seasonality, but I also added a few regressors to help it capture day-to-day changes in digital engagement. Specifically, I used five regressors:\n\n- n_games_total: total number of games played across the league that day (more games usually means more highlights and attention).\n\n- n_transactions_total: number of roster transactions that day (trades/call-ups, etc.), which can create news and discussion.\n\n- avg_win_pct: the average team winning percentage across the league, as a simple signal for overall team performance/competitiveness.\n\n- n_win_streak_teams: number of teams currently on a winning streak, since streaks often drive storylines and fan interest.\n\n- is_weekend: whether the day falls on a weekend, since fans may have more time to engage on Saturdays and Sundays.\n\nThe goal of these regressors is to give the model extra context about what’s happening in the league on a given day, beyond just the date itself.","metadata":{}},{"cell_type":"code","source":"def count_list_json(cell):\n    if pd.isna(cell) or cell == \"\":\n        return 0\n    try:\n        return len(json.loads(cell))\n    except Exception:\n        return 0\n\ndf[\"n_games_total\"] = df[\"games\"].apply(count_list_json)\ndf[\"n_transactions_total\"] = df[\"transactions\"].apply(count_list_json)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-18T04:28:26.803663Z","iopub.execute_input":"2025-12-18T04:28:26.804086Z","iopub.status.idle":"2025-12-18T04:28:27.015772Z","shell.execute_reply.started":"2025-12-18T04:28:26.804051Z","shell.execute_reply":"2025-12-18T04:28:27.014731Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def standings_summaries(cell):\n    # returns avg_win_pct, n_win_streak_teams\n    if pd.isna(cell) or cell == \"\":\n        return pd.Series([np.nan, 0])\n\n    try:\n        arr = json.loads(cell)\n        if not arr:\n            return pd.Series([np.nan, 0])\n\n        # pct is often a string like \"0.567\" or numeric; handle both\n        pcts = []\n        win_streak = 0\n        for d in arr:\n            pct = d.get(\"pct\", None)\n            if pct is not None and pct != \"\":\n                try:\n                    pcts.append(float(pct))\n                except:\n                    pass\n\n            streak = d.get(\"streakCode\", \"\")\n            if isinstance(streak, str) and streak.startswith(\"W\"):\n                win_streak += 1\n\n        avg_win_pct = np.mean(pcts) if len(pcts) else np.nan\n        return pd.Series([avg_win_pct, win_streak])\n\n    except Exception:\n        return pd.Series([np.nan, 0])\n\ndf[[\"avg_win_pct\",\"n_win_streak_teams\"]] = df[\"standings\"].apply(standings_summaries)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-18T04:28:27.215899Z","iopub.execute_input":"2025-12-18T04:28:27.216308Z","iopub.status.idle":"2025-12-18T04:28:27.534564Z","shell.execute_reply.started":"2025-12-18T04:28:27.216273Z","shell.execute_reply":"2025-12-18T04:28:27.533541Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# simple calendar regressors \ndf[\"is_weekend\"] = (df[\"ds\"].dt.dayofweek >= 5).astype(int)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-18T04:28:40.083931Z","iopub.execute_input":"2025-12-18T04:28:40.084317Z","iopub.status.idle":"2025-12-18T04:28:40.094481Z","shell.execute_reply.started":"2025-12-18T04:28:40.084285Z","shell.execute_reply":"2025-12-18T04:28:40.093533Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df[\"y\"] = (df[\"y1\"] + df[\"y2\"] + df[\"y3\"] + df[\"y4\"])/4","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-18T04:28:47.864061Z","iopub.execute_input":"2025-12-18T04:28:47.864438Z","iopub.status.idle":"2025-12-18T04:28:47.871222Z","shell.execute_reply.started":"2025-12-18T04:28:47.864406Z","shell.execute_reply":"2025-12-18T04:28:47.870166Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-18T04:28:52.704485Z","iopub.execute_input":"2025-12-18T04:28:52.705025Z","iopub.status.idle":"2025-12-18T04:28:52.747644Z","shell.execute_reply.started":"2025-12-18T04:28:52.704985Z","shell.execute_reply":"2025-12-18T04:28:52.746686Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_prophet = df[[\n    \"ds\",\"y\",\n    \"n_games_total\",\"n_transactions_total\",\n    \"avg_win_pct\",\"n_win_streak_teams\",\n    \"is_weekend\"\n]].copy()\n\ndf_prophet = df_prophet.dropna(subset=[\"y\"]).sort_values(\"ds\")\ndf_prophet[\"avg_win_pct\"] = df_prophet[\"avg_win_pct\"].fillna(df_prophet[\"avg_win_pct\"].median())\n\nfor c in [\"n_games_total\",\"n_transactions_total\",\"n_win_streak_teams\",\"is_weekend\"]:\n    df_prophet[c] = df_prophet[c].fillna(0)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-18T04:29:19.877550Z","iopub.execute_input":"2025-12-18T04:29:19.878534Z","iopub.status.idle":"2025-12-18T04:29:19.894393Z","shell.execute_reply.started":"2025-12-18T04:29:19.878492Z","shell.execute_reply":"2025-12-18T04:29:19.893250Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# EDA","metadata":{}},{"cell_type":"code","source":"color_pal = sns.color_palette()\ndf_prophet_plot = df[[\"ds\", \"y\"]]\ndf_prophet_plot[\"engagement_day\"] = df_prophet_plot[\"ds\"] + pd.Timedelta(days=1)\ndf_prophet_plot = df_prophet_plot[[\"engagement_day\", \"y\"]].set_index(\"engagement_day\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-18T04:29:36.192226Z","iopub.execute_input":"2025-12-18T04:29:36.193295Z","iopub.status.idle":"2025-12-18T04:29:36.204232Z","shell.execute_reply.started":"2025-12-18T04:29:36.193246Z","shell.execute_reply":"2025-12-18T04:29:36.203153Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"I started with a scatter plot of the full dataset to get a high-level view of the overall trend and day to day variability in MLB digital fan engagement. Each dot represents the daily league-wide engagement value. The red line shows the 7-day rolling average, which smooths out short term noise by averaging engagement over the past week, making the underlying pattern easier to see.\n\nAcross 2018–2021 (the time period covered by this dataset), engagement shows a clear seasonal cycle: it ramps up in late March and often peaks around April–May. My interpretation is that this lines up with the start of the MLB regular season because once the regular season begins, there are more highlights, storylines, standings movement, streaks, roster news, etc, so there’s simply more for fans to react to and engage with compared to the offseason or preseason.","metadata":{}},{"cell_type":"code","source":"import matplotlib.dates as mdates\nimport matplotlib.patheffects as pe\nfrom matplotlib.offsetbox import OffsetImage, AnnotationBbox\nimport matplotlib.image as mpimg\n\nMLB_NAVY = \"#041E42\"\nMLB_RED  = \"#BF0D3E\"\n\ny_col = \"y\"\ns = df_prophet_plot[y_col].sort_index()\nroll7 = s.rolling(7, min_periods=1).mean()\n\nfig, ax = plt.subplots(figsize=(13, 5.5), dpi=140)\nax.set_facecolor(\"#fbfbfd\")\nfig.patch.set_facecolor(\"white\")\n\n# season shading — but clip to data range later\nyears = pd.date_range(s.index.min().normalize(), s.index.max().normalize(), freq=\"YS\")\nfor y in years:\n    season_start = pd.Timestamp(year=y.year, month=3, day=1)\n    season_end   = pd.Timestamp(year=y.year, month=10, day=31)\n    ax.axvspan(season_start, season_end, color=MLB_NAVY, alpha=0.05, lw=0)\n\n# Raw points\nax.scatter(s.index, s.values, s=10, alpha=0.5, color=MLB_NAVY, edgecolors=\"none\", label=\"Daily\")\n\n# Rolling avg line (with white outline)\nline, = ax.plot(roll7.index, roll7.values, linewidth=2.8, color=MLB_RED, label=\"7-day rolling avg\")\nline.set_path_effects([pe.Stroke(linewidth=4.5, foreground=\"white\"), pe.Normal()])\n\n# Title\nax.set_title(\"MLB Digital Engagement — Daily Avg (mean of target1–target4)\", fontsize=13, pad=10)\n\nax.set_xlabel(\"Date\")\nax.set_ylabel(\"Engagement\")\n\nax.xaxis.set_major_locator(mdates.MonthLocator(interval=3))\nax.xaxis.set_major_formatter(mdates.DateFormatter(\"%b %Y\"))\nax.grid(True, alpha=0.25)\nax.spines[\"top\"].set_visible(False)\nax.spines[\"right\"].set_visible(False)\n\n# remove empty space: hard set x-limits to actual data\nax.set_xlim(s.index.min(), s.index.max())\n\n# tighter: move MLB text right\nax.text(0.925, 0.03, \"MLB\", transform=ax.transAxes,\n        ha=\"right\", va=\"bottom\", fontsize=16,\n        color=MLB_NAVY, alpha=0.85, weight=\"bold\", zorder=10)\n# Add logo to the right of the text, small and low\nlogo_path = \"/kaggle/input/mlb-logo/Major_League_Baseball_logo.svg.png\"\nlogo = mpimg.imread(logo_path)\n\n# move logo a bit left so it sits right next to the text\nimagebox = OffsetImage(logo, zoom=0.045)\nab = AnnotationBbox(\n    imagebox,\n    (0.928, 0.055),   # <-- was 0.935; move left toward the text\n    xycoords=\"axes fraction\",\n    frameon=False,\n    box_alignment=(0, 0.5),\n    pad=0,\n    zorder=10\n)\nax.add_artist(ab)\n\nax.legend(frameon=True, loc=\"upper left\")\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-18T04:34:22.102951Z","iopub.execute_input":"2025-12-18T04:34:22.103444Z","iopub.status.idle":"2025-12-18T04:34:22.654396Z","shell.execute_reply.started":"2025-12-18T04:34:22.103407Z","shell.execute_reply":"2025-12-18T04:34:22.653216Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Time Series Features\nI also wanted to see how engagement varies by day of week and season, so I first extracted time-based features from the datetime column (weekday, month, year, season, etc.) and then made a box plot to compare the distributions. From this plot, engagement is generally higher in spring and summer (blue and orange), which matches the in-season period when there are more games and storylines. Across the week, engagement is fairly consistent overall, but it tends to be slightly higher on Sundays.","metadata":{}},{"cell_type":"code","source":"from pandas.api.types import CategoricalDtype\n\ncat_type = CategoricalDtype(categories=['Monday','Tuesday',\n                                        'Wednesday',\n                                        'Thursday','Friday',\n                                        'Saturday','Sunday'],\n                            ordered=True)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-18T04:35:17.129301Z","iopub.execute_input":"2025-12-18T04:35:17.130339Z","iopub.status.idle":"2025-12-18T04:35:17.135859Z","shell.execute_reply.started":"2025-12-18T04:35:17.130295Z","shell.execute_reply":"2025-12-18T04:35:17.134873Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_prophet_time_series = df_prophet.copy()\ndf_prophet_time_series[\"engagement_day\"] = df_prophet_time_series[\"ds\"] + pd.Timedelta(days=1)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-18T04:35:23.697527Z","iopub.execute_input":"2025-12-18T04:35:23.697951Z","iopub.status.idle":"2025-12-18T04:35:23.705227Z","shell.execute_reply.started":"2025-12-18T04:35:23.697910Z","shell.execute_reply":"2025-12-18T04:35:23.704381Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def create_features(df, label=None):\n    \"\"\"\n    Creates time series features from datetime index.\n    \"\"\"\n    df = df.copy()\n    df['dayofweek'] = df['engagement_day'].dt.dayofweek\n    df['weekday'] = df['engagement_day'].dt.day_name()\n    df['weekday'] = df['weekday'].astype(cat_type)\n    df['quarter'] = df['engagement_day'].dt.quarter\n    df['month'] = df['engagement_day'].dt.month\n    df['year'] = df['engagement_day'].dt.year\n    df['dayofyear'] = df['engagement_day'].dt.dayofyear\n    df['dayofmonth'] = df['engagement_day'].dt.day\n    # df['weekofyear'] = df['ds'].dt.weekofyear\n    df['date_offset'] = (df['engagement_day'].dt.month*100 + df['engagement_day'].dt.day - 320)%1300\n\n    df['season'] = pd.cut(df['date_offset'], [0, 300, 602, 900, 1300], \n                          labels=['Spring', 'Summer', 'Fall', 'Winter']\n                   )\n    X = df[['dayofweek','quarter','month','year',\n           'dayofyear','dayofmonth','weekday',\n           'season']]\n    if label:\n        y = df[label]\n        return X, y\n    return X\n\nX, y = create_features(df_prophet_time_series, label='y')\nfeatures_and_target = pd.concat([X, y], axis=1)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-18T04:35:33.953657Z","iopub.execute_input":"2025-12-18T04:35:33.954079Z","iopub.status.idle":"2025-12-18T04:35:33.982850Z","shell.execute_reply.started":"2025-12-18T04:35:33.954043Z","shell.execute_reply":"2025-12-18T04:35:33.981583Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(10, 5))\nsns.boxplot(data=features_and_target.dropna(),\n            x='weekday',\n            y='y',\n            hue='season',\n            ax=ax,\n            linewidth=1)\nax.set_title(\"Weekday vs MLB Engagement (by Season)\")\nax.set_xlabel(\"Weekday\")\nax.set_ylabel(\"Next-day Avg Engagement\")\nax.legend(bbox_to_anchor=(1, 1))\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-18T04:35:41.090816Z","iopub.execute_input":"2025-12-18T04:35:41.091245Z","iopub.status.idle":"2025-12-18T04:35:41.731217Z","shell.execute_reply.started":"2025-12-18T04:35:41.091208Z","shell.execute_reply":"2025-12-18T04:35:41.730133Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Train/Test Split\nTo prepare the data for modeling and evaluation, I split the time series into training and testing sets (with the test set covering later, “future” dates). I also plotted a scatter chart to visually compare how engagement behaves in the train vs. test periods, which helps confirm whether the test window follows a similar distribution/pattern or includes noticeable shifts that could affect forecasting performance.","metadata":{}},{"cell_type":"code","source":"df_train = df_prophet.iloc[:-200, :]\ndf_test= df_prophet.iloc[-200:, :]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-18T04:36:29.457983Z","iopub.execute_input":"2025-12-18T04:36:29.458966Z","iopub.status.idle":"2025-12-18T04:36:29.464179Z","shell.execute_reply.started":"2025-12-18T04:36:29.458925Z","shell.execute_reply":"2025-12-18T04:36:29.463220Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Plot to show train vs test engagement\ndf_train_plot = df_train[[\"ds\", \"y\"]].copy()\ndf_test_plot = df_test[[\"ds\", \"y\"]].copy()\ndf_train_plot[\"engagement_day\"] = df_train_plot[\"ds\"] + pd.Timedelta(days=1)\ndf_test_plot[\"engagement_day\"] = df_test_plot[\"ds\"] + pd.Timedelta(days=1)\n\ndf_train_plot = df_train_plot[[\"engagement_day\", \"y\"]].set_index(\"engagement_day\")\ndf_test_plot = df_test_plot[[\"engagement_day\", \"y\"]].set_index(\"engagement_day\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-18T04:36:54.590717Z","iopub.execute_input":"2025-12-18T04:36:54.593800Z","iopub.status.idle":"2025-12-18T04:36:54.623385Z","shell.execute_reply.started":"2025-12-18T04:36:54.593685Z","shell.execute_reply":"2025-12-18T04:36:54.622484Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"MLB_NAVY = \"#041E42\"\nMLB_RED  = \"#BF0D3E\"\nTEST_GOLD = \"#F5B700\"\n\ns_train = df_train_plot[\"y\"].sort_index()\ns_test  = df_test_plot[\"y\"].sort_index()\n\nr_train = s_train.rolling(7, min_periods=1).mean()\nr_test  = s_test.rolling(7, min_periods=1).mean()\n\nsplit_date = s_test.index.min()\n\nfig, ax = plt.subplots(figsize=(13, 5.5), dpi=160)\nax.set_facecolor(\"#fbfbfd\")\nfig.patch.set_facecolor(\"white\")\n\n# --- background shading for train/test ---\nax.axvspan(s_train.index.min(), split_date, color=MLB_NAVY, alpha=0.04, lw=0)\nax.axvspan(split_date, s_test.index.max(), color=TEST_GOLD, alpha=0.06, lw=0)\n\n# --- scatter (lighter) ---\nax.scatter(s_train.index, s_train.values, s=9,  alpha=0.18, color=MLB_NAVY, edgecolors=\"none\", label=\"Train (daily)\")\nax.scatter(s_test.index,  s_test.values,  s=10, alpha=0.22, color=TEST_GOLD, edgecolors=\"none\", label=\"Test (daily)\")\n\n# --- rolling lines (hero) with white outline ---\nl1, = ax.plot(r_train.index, r_train.values, color=MLB_RED, linewidth=2.8, label=\"Train (7-day avg)\")\nl1.set_path_effects([pe.Stroke(linewidth=4.6, foreground=\"white\"), pe.Normal()])\n\nl2, = ax.plot(r_test.index, r_test.values, color=\"black\", alpha=0.85, linewidth=2.8, label=\"Test (7-day avg)\")\nl2.set_path_effects([pe.Stroke(linewidth=4.6, foreground=\"white\"), pe.Normal()])\n\n# --- split line + label (cleaner positioning) ---\nax.axvline(split_date, linestyle=\"--\", linewidth=2, alpha=0.55, color=MLB_NAVY)\nax.text(split_date, 0.98, \" Test starts\", transform=ax.get_xaxis_transform(),\n        ha=\"left\", va=\"top\", fontsize=10, color=MLB_NAVY, alpha=0.85)\n\n# --- titles/labels ---\nax.set_title(\"MLB Next-day Engagement — Train vs Test\", fontsize=14, pad=10)\nax.set_xlabel(\"Engagement day\")\nax.set_ylabel(\"Avg engagement (mean of target1–target4)\")\n\n# --- date formatting ---\nax.xaxis.set_major_locator(mdates.MonthLocator(interval=3))\nax.xaxis.set_major_formatter(mdates.DateFormatter(\"%b %Y\"))\n\n# --- grid + spines ---\nax.grid(True, alpha=0.22)\nax.spines[\"top\"].set_visible(False)\nax.spines[\"right\"].set_visible(False)\n\n# --- limits so no extra whitespace ---\nax.set_xlim(s_train.index.min(), s_test.index.max())\n\n# --- nicer legend ---\nleg = ax.legend(frameon=True, loc=\"upper left\")\nleg.get_frame().set_alpha(0.95)\nleg.get_frame().set_edgecolor(\"#dddddd\")\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-18T04:36:56.481806Z","iopub.execute_input":"2025-12-18T04:36:56.482591Z","iopub.status.idle":"2025-12-18T04:36:57.023097Z","shell.execute_reply.started":"2025-12-18T04:36:56.482548Z","shell.execute_reply":"2025-12-18T04:36:57.021947Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Prophet Model\nThe main reason I decided to use Facebook/Meta’s Prophet time series forecasting model is because MLB fan engagement is heavily time-driven. Engagement isn’t random — it follows patterns tied to the MLB calendar, like the season ramping up around March/April, plus predictable weekly effects (weekends) and holiday/event effects when fans have more time and there’s more content to react to. Prophet is built to model these kinds of trend + seasonality + holiday effects in a way that’s interpretable, which makes it really useful for an analytics setting.\n\nEven though lot of deep learning approaches can get strong scores, they don’t always explicitly capture or explain these time-based fluctuations unless you engineer them carefully, and they’re harder to communicate to non-technical stakeholders. With Prophet, I can both forecast engagement and clearly show what’s driving it (weekly seasonality, yearly seasonality, and special dates), which is exactly the kind of insight I want to demonstrate.","metadata":{}},{"cell_type":"code","source":"df_train[\"ds\"] = df_train[\"ds\"] + pd.Timedelta(days=1)\ndf_test[\"ds\"] = df_test[\"ds\"] + pd.Timedelta(days=1)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-18T04:37:34.341393Z","iopub.execute_input":"2025-12-18T04:37:34.342667Z","iopub.status.idle":"2025-12-18T04:37:34.351588Z","shell.execute_reply.started":"2025-12-18T04:37:34.342618Z","shell.execute_reply":"2025-12-18T04:37:34.350384Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"m = Prophet(\n    weekly_seasonality=True,\n    yearly_seasonality=True,\n    daily_seasonality=False,\n    interval_width=0.9\n)\n\n# US holidays \nm.add_country_holidays(country_name=\"US\")\n\n# Add regressors\nm.add_regressor(\"n_games_total\")\nm.add_regressor(\"n_transactions_total\")\nm.add_regressor(\"avg_win_pct\")\nm.add_regressor(\"n_win_streak_teams\")\nm.add_regressor(\"is_weekend\")\n\nm.fit(df_train)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-18T04:37:45.410597Z","iopub.execute_input":"2025-12-18T04:37:45.411362Z","iopub.status.idle":"2025-12-18T04:37:48.830697Z","shell.execute_reply.started":"2025-12-18T04:37:45.411323Z","shell.execute_reply":"2025-12-18T04:37:48.829569Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test_fcst = m.predict(df_test)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-18T04:37:56.012992Z","iopub.execute_input":"2025-12-18T04:37:56.013656Z","iopub.status.idle":"2025-12-18T04:37:56.185394Z","shell.execute_reply.started":"2025-12-18T04:37:56.013610Z","shell.execute_reply":"2025-12-18T04:37:56.184313Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The blue section shows Prophet’s forecasted engagement over the test period. The darker blue line is the model’s point forecast, and the lighter blue band represents the model’s prediction interval (uncertainty range) around that forecast for each date.","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(10, 5))\nfig = m.plot(test_fcst, ax=ax)\nax.set_title('Prophet Forecast')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-18T04:38:11.372182Z","iopub.execute_input":"2025-12-18T04:38:11.372567Z","iopub.status.idle":"2025-12-18T04:38:11.623421Z","shell.execute_reply.started":"2025-12-18T04:38:11.372535Z","shell.execute_reply":"2025-12-18T04:38:11.622281Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"These plots below break down the different Prophet components and show what the model learned from the data. From the trend component, Prophet picks up an overall decrease in engagement over the full time period, which matches the general direction we see in the dataset. The holiday component suggests engagement tends to spike around major holidays like Thanksgiving and Christmas. The weekly component captures a weekend effect, where engagement is generally higher on Saturdays and Sundays and lower during weekdays. Finally, the yearly component captures the strong seasonal pattern where engagement ramps up around March–April, which lines up with the start of the MLB regular season when fan attention typically increases.","metadata":{}},{"cell_type":"code","source":"fig = m.plot_components(test_fcst)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-18T04:38:45.591371Z","iopub.execute_input":"2025-12-18T04:38:45.592241Z","iopub.status.idle":"2025-12-18T04:38:46.693314Z","shell.execute_reply.started":"2025-12-18T04:38:45.592181Z","shell.execute_reply":"2025-12-18T04:38:46.692324Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Compare forecast with actuals\nThis plot compares the actual engagement values in the test period (yellow dots) with Prophet’s forecast (red line) and its prediction interval (light blue band). Overall, the model does capture the main seasonal pattern, especially the ramp up into March and April, which is consistent with the start of the regular MLB season that leads to increase in fan attention.\n\nHowever, there’s still noticeable deviation between the actual and predicted values. In this test window, the forecast is often higher than the actual engagement, which suggests the model may be overestimating the level of engagement. A few realistic reasons for this are that engagement can shift due to real-world events that are not fully explained by the time features and regressors (for example, COVID-related changes in fan behavior), and that league wide engagement is also driven by sudden spikes from major news days (big trades, injuries, viral moments) that are hard to predict from aggregate signals. Another factor is that we are forecasting a noisy daily metric, so some mismatch is expected even when the overall trend is right. ","metadata":{}},{"cell_type":"code","source":"\nMLB_NAVY = \"#041E42\"\nMLB_RED  = \"#BF0D3E\"\nTEST_GOLD = \"#F5B700\"\n\n# df_test must have ds, y\n# test_fcst must have ds, yhat, yhat_lower, yhat_upper\n# df_train optional for context (ds,y)\n\nfig, ax = plt.subplots(figsize=(13, 5.5), dpi=160)\nax.set_facecolor(\"#fbfbfd\")\nfig.patch.set_facecolor(\"white\")\n\n# 1) history (context)\nif \"df_train\" in globals():\n    ax.scatter(df_train[\"ds\"], df_train[\"y\"],\n               s=10, alpha=0.12, color=MLB_NAVY, edgecolors=\"none\",\n               label=\"Train (actual)\")\n\n# 2) uncertainty band (subtle)\nax.fill_between(test_fcst[\"ds\"],\n                test_fcst[\"yhat_lower\"],\n                test_fcst[\"yhat_upper\"],\n                alpha=0.18, linewidth=0,\n                label=\"Forecast interval\")\n\n# 3) forecast line (hero)\nline, = ax.plot(test_fcst[\"ds\"], test_fcst[\"yhat\"],\n                color=MLB_RED, linewidth=2.8,\n                label=\"Prophet forecast (yhat)\")\nline.set_path_effects([pe.Stroke(linewidth=4.6, foreground=\"white\"), pe.Normal()])\n\n# 4) actual test points (nice + readable)\nax.scatter(df_test[\"ds\"], df_test[\"y\"],\n           s=22, alpha=0.75, color=TEST_GOLD,\n           edgecolors=\"white\", linewidth=0.6,\n           label=\"Test (actual)\")\n\n# 5) split line\nsplit_date = df_test[\"ds\"].min()\nax.axvline(split_date, linestyle=\"--\", linewidth=2, alpha=0.55, color=MLB_NAVY)\nax.text(split_date, 0.98, \" Test starts\", transform=ax.get_xaxis_transform(),\n        ha=\"left\", va=\"top\", fontsize=10, color=MLB_NAVY, alpha=0.85)\n\n# labels\nax.set_title(\"Prophet Forecast vs Actuals — MLB Next-day Engagement\", fontsize=14, pad=10)\nax.set_xlabel(\"Date\")\nax.set_ylabel(\"Avg engagement (mean of target1–target4)\")\n\n# dates/grid/spines\nax.xaxis.set_major_locator(mdates.MonthLocator(interval=3))\nax.xaxis.set_major_formatter(mdates.DateFormatter(\"%b %Y\"))\nax.grid(True, alpha=0.22)\nax.spines[\"top\"].set_visible(False)\nax.spines[\"right\"].set_visible(False)\n\n# x-limits to remove extra whitespace\nax.set_xlim(min(df_train[\"ds\"].min(), df_test[\"ds\"].min()) if \"df_train\" in globals() else df_test[\"ds\"].min(),\n            max(df_test[\"ds\"].max(), test_fcst[\"ds\"].max()))\n\n# legend (clean)\nleg = ax.legend(frameon=True, loc=\"upper left\")\nleg.get_frame().set_alpha(0.95)\nleg.get_frame().set_edgecolor(\"#dddddd\")\n\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-18T04:40:01.924686Z","iopub.execute_input":"2025-12-18T04:40:01.926263Z","iopub.status.idle":"2025-12-18T04:40:02.496076Z","shell.execute_reply.started":"2025-12-18T04:40:01.926197Z","shell.execute_reply":"2025-12-18T04:40:02.494810Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Model Evaluation\nI evaluated the model using a simple seasonal baseline where the predicted daily fan engagement is just the engagement from the same calendar day last year. This baseline is useful because it captures basic yearly seasonality, so it helps check whether Prophet is actually learning additional structure like trend shifts, weekly patterns, and holiday effects instead of just repeating last year.\n\nI then compared both the baseline and Prophet using MAE and RMSE. In my results, Prophet achieves noticeably lower MAE and RMSE than the seasonal naive baseline, suggesting it’s doing more than copying last year’s engagement and is better at tracking the short-term fluctuations in engagement.","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import mean_absolute_error, mean_squared_error\n\ndef rmse(y_true, y_pred):\n    return np.sqrt(mean_squared_error(y_true, y_pred))\n\n# 1) Prophet eval\nmerged_prophet = df_test.merge(test_fcst[[\"ds\",\"yhat\"]], on=\"ds\", how=\"inner\").dropna()\nprint(\"Prophet MAE:\", mean_absolute_error(merged_prophet[\"y\"], merged_prophet[\"yhat\"]))\nprint(\"Prophet RMSE:\", rmse(merged_prophet[\"y\"], merged_prophet[\"yhat\"]))\n\n# 2) Seasonal naive - same day last year\n# Build a lookup table from history (train) keyed by date\nhistory = df_train[[\"ds\",\"y\"]].copy()\nhistory[\"ds\"] = pd.to_datetime(history[\"ds\"])\nhistory = history.sort_values(\"ds\")\n\ntest = df_test[[\"ds\",\"y\"]].copy()\ntest[\"ds\"] = pd.to_datetime(test[\"ds\"])\ntest = test.sort_values(\"ds\")\n\n# For each test day, look up y from exactly 365 days earlier\ntest[\"ds_last_year\"] = test[\"ds\"] - pd.Timedelta(days=365)\n\nlookup = history.rename(columns={\"ds\": \"ds_last_year\", \"y\": \"y_last_year\"})\nseasonal = test.merge(lookup, on=\"ds_last_year\", how=\"left\").dropna(subset=[\"y_last_year\"])\n\nprint(\"Seasonal naive (t-365d) MAE:\", mean_absolute_error(seasonal[\"y\"], seasonal[\"y_last_year\"]))\nprint(\"Seasonal naive (t-365d) RMSE:\", rmse(seasonal[\"y\"], seasonal[\"y_last_year\"]))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-18T04:40:49.072558Z","iopub.execute_input":"2025-12-18T04:40:49.072991Z","iopub.status.idle":"2025-12-18T04:40:49.692380Z","shell.execute_reply.started":"2025-12-18T04:40:49.072955Z","shell.execute_reply":"2025-12-18T04:40:49.691315Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Decision Dashboard: Insights Obtained\n\nSince the model does a decent job forecasting engagement, we can use the test set forecast to get more actionable insights. In the interactive plot below, I convert the predicted engagement into an action level (low/medium/high) based on where the forecast falls relative to historical engagement levels, where higher percentiles are labeled as “high”. When the model predicts unusually high engagement, I treat it as a high-action day (marked with a purple circle), meaning teams could plan ahead by scheduling their best content, timing campaigns, and making sure there’s enough coverage. I also include the prediction interval to reflect uncertainty, since the exact engagement level can still vary day to day.","metadata":{}},{"cell_type":"code","source":" import plotly.graph_objects as go\n\n# thresholds from training distribution\np75 = df_train[\"y\"].quantile(0.75)\np90 = df_train[\"y\"].quantile(0.90)\n\nplot = df_test.merge(test_fcst[[\"ds\",\"yhat\",\"yhat_lower\",\"yhat_upper\"]], on=\"ds\", how=\"inner\").dropna()\nplot = plot.sort_values(\"ds\")\n\n# action levels\nplot[\"action\"] = np.where(plot[\"yhat\"] >= p90, \"High\",\n                  np.where(plot[\"yhat\"] >= p75, \"Medium\", \"Low\"))\n\nfig = go.Figure()\n\n# uncertainty band\nfig.add_trace(go.Scatter(\n    x=pd.concat([plot[\"ds\"], plot[\"ds\"][::-1]]),\n    y=pd.concat([plot[\"yhat_upper\"], plot[\"yhat_lower\"][::-1]]),\n    fill=\"toself\",\n    line=dict(width=0),\n    name=\"Forecast interval\",\n    hoverinfo=\"skip\",\n    opacity=0.2\n))\n\n# forecast line\nfig.add_trace(go.Scatter(\n    x=plot[\"ds\"], y=plot[\"yhat\"],\n    mode=\"lines\",\n    name=\"Prophet forecast\",\n))\n\n# actual points\nfig.add_trace(go.Scatter(\n    x=plot[\"ds\"], y=plot[\"y\"],\n    mode=\"markers\",\n    name=\"Actual\",\n    marker=dict(size=6),\n))\n\n# highlight action-level points (optional markers on top)\nfor level in [\"High\",\"Medium\",\"Low\"]:\n    sub = plot[plot[\"action\"] == level]\n    fig.add_trace(go.Scatter(\n        x=sub[\"ds\"], y=sub[\"yhat\"],\n        mode=\"markers\",\n        name=f\"Action: {level}\",\n        marker=dict(size=7, symbol=\"circle-open\"),\n        hovertemplate=\"Date=%{x}<br>yhat=%{y:.3f}<extra></extra>\"\n    ))\n\nfig.update_layout(\n    title=\"Decision View: Forecast + Uncertainty + Action Level\",\n    xaxis_title=\"Date\",\n    yaxis_title=\"Engagement\",\n    legend_title=\"\",\n    template=\"plotly_white\",\n    height=520\n)\n\n# fig.show()\niplot(fig)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-18T06:13:46.457434Z","iopub.execute_input":"2025-12-18T06:13:46.458828Z","iopub.status.idle":"2025-12-18T06:13:46.551137Z","shell.execute_reply.started":"2025-12-18T06:13:46.458740Z","shell.execute_reply":"2025-12-18T06:13:46.549667Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"I also wanted to look at the model’s top 10 biggest mistakes, based on the absolute error between the actual engagement y and the predicted engagement yhat. From these results, I noticed the model underpredicted engagement on 11/27, which is right after Thanksgiving, when people are still off work, and generally more active. The model also tends to underpredict engagement on days in mid to late March, when MLB is approaching the regular season and fan attention starts ramping up faster than the model expects.\n\nOn the other hand, the model overpredicts engagement during mid to late January and early February, which is still offseason. Even when there are a lot of transactions happening, overall fan engagement can stay low compared to what the model predicts. For example, on 1/16/2021 there were 429 transactions, but actual engagement was still much lower than the forecast, likely because there are no regular season games and less daily content to consistently drive attention.","metadata":{}},{"cell_type":"code","source":"# Merge actual vs predicted\nerr_df = (\n    df_test\n    .merge(test_fcst[[\"ds\", \"yhat\"]], on=\"ds\", how=\"inner\")\n    .dropna()\n    .sort_values(\"ds\")\n)\n\n# Compute errors\nerr_df[\"error\"] = err_df[\"y\"] - err_df[\"yhat\"]          # positive = underpredicted\nerr_df[\"abs_error\"] = err_df[\"error\"].abs()\n\n# Top 10 biggest misses\ntop10 = (\n    err_df.sort_values(\"abs_error\", ascending=False)\n    .head(10)\n    .copy()\n)\n\n# Optional: nicer formatting\ntop10[\"ds\"] = pd.to_datetime(top10[\"ds\"]).dt.date\n# top10 = top10[[\"ds\", \"y\", \"yhat\", \"error\", \"abs_error\"]]\n\ntop10","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-18T06:43:48.383683Z","iopub.execute_input":"2025-12-18T06:43:48.384106Z","iopub.status.idle":"2025-12-18T06:43:48.410883Z","shell.execute_reply.started":"2025-12-18T06:43:48.384063Z","shell.execute_reply":"2025-12-18T06:43:48.409550Z"}},"outputs":[],"execution_count":null}]}