{
  "id": 256658,
  "title": "9th place solution",
  "url": "/competitions/mlb-player-digital-engagement-forecasting/discussion/256658",
  "author_name": "Mark Tenenholtz",
  "post_date": "2021-08-02T16:44:04.220000",
  "votes": 38,
  "comment_count": 9,
  "views": 0,
  "content": "<p>This was one heck of a competition. There was no shortage of new things to try at any point given the extent of the data. For that, I'd like to thank both the MLB and Kaggle.</p>\n<h2>High-level strategy</h2>\n<p>It's no secret that lagged features were very important for this competition. I found this, too, and for a while my solutions were heavily reliant on lagged target variables. I suspect most high scoring submissions were, as well. However, there was a lot of uncertainty about what the gap between the ground truth and the inference period would be. I had considered doing some augmentations for robustness, like dropping out the n most recent days of target lags randomly, but this just killed my model's performance. So, I ended up abandoning this strategy. While I do have some target feature aggregations in my solution, as far as lags go, my final submission relies <strong>only on lags of known features</strong> like box score, standings, and games features.</p>\n<h2>Models used</h2>\n<p>I tried out a ton of different models that didn't work. Here's a quick summary:</p>\n<p>What didn't work at all:</p>\n<ul>\n<li>Denoising autoencoders</li>\n</ul>\n<p>What worked decently, but not well enough:</p>\n<ul>\n<li>LightGBM, CatBoost</li>\n<li>1D CNN (from the <a href=\"https://www.kaggle.com/c/lish-moa/discussion/202256\" target=\"_blank\">MoA competition</a>)</li>\n<li>Time series models on lagged targets (GRU, Transformer, DeepAR, Temporal Fusion Transformer)</li>\n</ul>\n<p>My final submission ended up being an ensemble of a GRU and a Transformer that rely on non-target lags. The GRU was my best model, achieving ~1.25 on its own. The Transformer scored about ~1.27 on its own, but it proved to be very useful in an ensemble.</p>\n<h2>Validation</h2>\n<p>I did a simple time series split. Before the training data was updated, I used July and August in 2020, and then April in 2021. After the updated data was released, I used August 2020 and June and July in 2021.</p>\n<h2>Features used</h2>\n<p>Below is a list of all the types of features that I used. All lagged features go back 2 weeks.</p>\n<ul>\n<li>Mean/median/std/min/max of each target from the previous month</li>\n<li>Embeddings of playerId, teamId, player position, and player status</li>\n<li>Player box score lags<ul>\n<li>'gamesPlayedBatting', 'hits', 'doubles', 'triples', 'runsScored',<br>\n'homeRuns', 'hitByPitch', 'totalBases', 'rbi', 'stolenBases', 'assists', <br>\n'gamesPlayedPitching', 'completeGamesPitching', 'shutoutsPitching', 'earnedRuns', 'winsPitching', 'strikeOutsPitching', 'hitsPitching', 'saveOpportunities', 'saves', 'holds', 'inningsPitched'</li></ul></li>\n<li>Lagged game features<ul>\n<li>'gameTimeUTC', 'wasSigned', 'wasTraded', 'teamWins', 'teamLosses', 'teamScore', 'isHome', 'teamWon', 'scoreDiff'</li>\n<li>The game time feature only includes the hour the game was played. The idea here was a lot of the digital engagement resulting from a, say, 1 PM game would manifest itself on the day the game was played, whereas the engagement from a night game would all happen on the next day.</li></ul></li>\n<li>Lagged standings features<ul>\n<li>'wins', 'losses', 'pct', 'xWinLossPct', 'divisionRank', 'lastTenWins', 'lastTenLosses'</li></ul></li>\n<li>Lagged cumulative features (all box score features summed up by season)</li>\n<li>Lagged transactional features, which were just a flag indicating if the player had either been traded or released on the particular day.</li>\n<li>Player followers, team followers</li>\n</ul>\n<p>I tried to use features that exploited the scaling of the data, such as the figuring out the minimum increment between target features for a player and rescaling their target features to the non-scaled \"actual\" feature, but it didn't really help my model at all. I was pretty surprised by this.</p>\n<h2>Final submissions</h2>\n<p>I noticed that including as much data as possible helped immensely, so I knew that I wanted to make at least one of my final solutions have models that were trained on 100% of the available data. However, my training curves for both the Transformer and GRU were <a href=\"https://drive.google.com/file/d/1OR2vXYhGv8W3VoDWyLDicFef79UTXjs-/view?usp=sharing\" target=\"_blank\">quite bumpy</a>. So, I took checkpoints every 200 epochs starting from step 1500 while training blind on 100% of the data. I also ensembled across several seeds.</p>\n<p>To be safe, my second submission used models that used the last 30 days of data available as validation data. This is really just a hedge, but I'm interested to see how it does.</p>\n<h2>Other notes</h2>\n<ul>\n<li>All models were trained in PyTorch using PyTorch Lightning</li>\n<li>Editor was VSCode and I made heavy use of the TensorBoard integration</li>\n<li><a href=\"https://www.kaggle.com/nyanpn\" target=\"_blank\">@nyanpn</a>'s submission emulator was invaluable near the end of the competition</li>\n<li>I used Optuna for hyperparameter tuning.</li>\n<li>To reduce submission bugs, I used a dataset that has example versions of roster data, game data, box score data, etc. This allowed me to have a \"default\" version of the data to return when it is null.</li>\n<li>As some other competitors noted, the predictions from even my best models look somewhat suspect. I was thrilled whenever I saw my model nail an upward spike in engagement. However, I think the MLB is acutely aware that some major variations in engagement happen because of exogenous factors like news stories. Given the MLB's fairly recent expansion of their social media endeavors, it makes sense that they would want to focus on purely the game-level factors that drive engagement, as this is all they really can act on. Maybe this is a faulty argument, but because of this I think MAE actually made pretty decent sense as a competition metric -- MSE would penalize you too much for factors that are completely unpredictable from game-level data.</li>\n</ul>\n<p>If you made it this far, thanks for reading! I had a lot of fun during this competition and can't wait to see how things shake out.</p>\n<p>EDIT: I've made my inference notebook public here: <a href=\"https://www.kaggle.com/marktenenholtz/mlb-rnn-transformer-final-public\" target=\"_blank\">https://www.kaggle.com/marktenenholtz/mlb-rnn-transformer-final-public</a></p>\n<p>A basic run-down of the models is that they're essentially the same, with the GRU and Transformer swapped in as different backbones. Other than that, they both take in the embedded categorical features as well as the lagged features, feed those into the backbone as a sequence, and then there's 3 convolutions that are applied to the output (for whatever reason, this worked better than a simple head). After the convolution, there's a set of linear layers that comprises the head of the model, and there's a unique head for each target. There are some minor differences from this in the Transformer model, but this is more or less how it works.</p>",
  "messages": [
    {
      "id": 1409239,
      "postDate": "2021-08-02T16:44:04.220Z",
      "content": "<p>This was one heck of a competition. There was no shortage of new things to try at any point given the extent of the data. For that, I'd like to thank both the MLB and Kaggle.</p>\n<h2>High-level strategy</h2>\n<p>It's no secret that lagged features were very important for this competition. I found this, too, and for a while my solutions were heavily reliant on lagged target variables. I suspect most high scoring submissions were, as well. However, there was a lot of uncertainty about what the gap between the ground truth and the inference period would be. I had considered doing some augmentations for robustness, like dropping out the n most recent days of target lags randomly, but this just killed my model's performance. So, I ended up abandoning this strategy. While I do have some target feature aggregations in my solution, as far as lags go, my final submission relies <strong>only on lags of known features</strong> like box score, standings, and games features.</p>\n<h2>Models used</h2>\n<p>I tried out a ton of different models that didn't work. Here's a quick summary:</p>\n<p>What didn't work at all:</p>\n<ul>\n<li>Denoising autoencoders</li>\n</ul>\n<p>What worked decently, but not well enough:</p>\n<ul>\n<li>LightGBM, CatBoost</li>\n<li>1D CNN (from the <a href=\"https://www.kaggle.com/c/lish-moa/discussion/202256\" target=\"_blank\">MoA competition</a>)</li>\n<li>Time series models on lagged targets (GRU, Transformer, DeepAR, Temporal Fusion Transformer)</li>\n</ul>\n<p>My final submission ended up being an ensemble of a GRU and a Transformer that rely on non-target lags. The GRU was my best model, achieving ~1.25 on its own. The Transformer scored about ~1.27 on its own, but it proved to be very useful in an ensemble.</p>\n<h2>Validation</h2>\n<p>I did a simple time series split. Before the training data was updated, I used July and August in 2020, and then April in 2021. After the updated data was released, I used August 2020 and June and July in 2021.</p>\n<h2>Features used</h2>\n<p>Below is a list of all the types of features that I used. All lagged features go back 2 weeks.</p>\n<ul>\n<li>Mean/median/std/min/max of each target from the previous month</li>\n<li>Embeddings of playerId, teamId, player position, and player status</li>\n<li>Player box score lags<ul>\n<li>'gamesPlayedBatting', 'hits', 'doubles', 'triples', 'runsScored',<br>\n'homeRuns', 'hitByPitch', 'totalBases', 'rbi', 'stolenBases', 'assists', <br>\n'gamesPlayedPitching', 'completeGamesPitching', 'shutoutsPitching', 'earnedRuns', 'winsPitching', 'strikeOutsPitching', 'hitsPitching', 'saveOpportunities', 'saves', 'holds', 'inningsPitched'</li></ul></li>\n<li>Lagged game features<ul>\n<li>'gameTimeUTC', 'wasSigned', 'wasTraded', 'teamWins', 'teamLosses', 'teamScore', 'isHome', 'teamWon', 'scoreDiff'</li>\n<li>The game time feature only includes the hour the game was played. The idea here was a lot of the digital engagement resulting from a, say, 1 PM game would manifest itself on the day the game was played, whereas the engagement from a night game would all happen on the next day.</li></ul></li>\n<li>Lagged standings features<ul>\n<li>'wins', 'losses', 'pct', 'xWinLossPct', 'divisionRank', 'lastTenWins', 'lastTenLosses'</li></ul></li>\n<li>Lagged cumulative features (all box score features summed up by season)</li>\n<li>Lagged transactional features, which were just a flag indicating if the player had either been traded or released on the particular day.</li>\n<li>Player followers, team followers</li>\n</ul>\n<p>I tried to use features that exploited the scaling of the data, such as the figuring out the minimum increment between target features for a player and rescaling their target features to the non-scaled \"actual\" feature, but it didn't really help my model at all. I was pretty surprised by this.</p>\n<h2>Final submissions</h2>\n<p>I noticed that including as much data as possible helped immensely, so I knew that I wanted to make at least one of my final solutions have models that were trained on 100% of the available data. However, my training curves for both the Transformer and GRU were <a href=\"https://drive.google.com/file/d/1OR2vXYhGv8W3VoDWyLDicFef79UTXjs-/view?usp=sharing\" target=\"_blank\">quite bumpy</a>. So, I took checkpoints every 200 epochs starting from step 1500 while training blind on 100% of the data. I also ensembled across several seeds.</p>\n<p>To be safe, my second submission used models that used the last 30 days of data available as validation data. This is really just a hedge, but I'm interested to see how it does.</p>\n<h2>Other notes</h2>\n<ul>\n<li>All models were trained in PyTorch using PyTorch Lightning</li>\n<li>Editor was VSCode and I made heavy use of the TensorBoard integration</li>\n<li><a href=\"https://www.kaggle.com/nyanpn\" target=\"_blank\">@nyanpn</a>'s submission emulator was invaluable near the end of the competition</li>\n<li>I used Optuna for hyperparameter tuning.</li>\n<li>To reduce submission bugs, I used a dataset that has example versions of roster data, game data, box score data, etc. This allowed me to have a \"default\" version of the data to return when it is null.</li>\n<li>As some other competitors noted, the predictions from even my best models look somewhat suspect. I was thrilled whenever I saw my model nail an upward spike in engagement. However, I think the MLB is acutely aware that some major variations in engagement happen because of exogenous factors like news stories. Given the MLB's fairly recent expansion of their social media endeavors, it makes sense that they would want to focus on purely the game-level factors that drive engagement, as this is all they really can act on. Maybe this is a faulty argument, but because of this I think MAE actually made pretty decent sense as a competition metric -- MSE would penalize you too much for factors that are completely unpredictable from game-level data.</li>\n</ul>\n<p>If you made it this far, thanks for reading! I had a lot of fun during this competition and can't wait to see how things shake out.</p>\n<p>EDIT: I've made my inference notebook public here: <a href=\"https://www.kaggle.com/marktenenholtz/mlb-rnn-transformer-final-public\" target=\"_blank\">https://www.kaggle.com/marktenenholtz/mlb-rnn-transformer-final-public</a></p>\n<p>A basic run-down of the models is that they're essentially the same, with the GRU and Transformer swapped in as different backbones. Other than that, they both take in the embedded categorical features as well as the lagged features, feed those into the backbone as a sequence, and then there's 3 convolutions that are applied to the output (for whatever reason, this worked better than a simple head). After the convolution, there's a set of linear layers that comprises the head of the model, and there's a unique head for each target. There are some minor differences from this in the Transformer model, but this is more or less how it works.</p>",
      "rawMarkdown": "This was one heck of a competition. There was no shortage of new things to try at any point given the extent of the data. For that, I'd like to thank both the MLB and Kaggle.\n\n## High-level strategy\n\nIt's no secret that lagged features were very important for this competition. I found this, too, and for a while my solutions were heavily reliant on lagged target variables. I suspect most high scoring submissions were, as well. However, there was a lot of uncertainty about what the gap between the ground truth and the inference period would be. I had considered doing some augmentations for robustness, like dropping out the n most recent days of target lags randomly, but this just killed my model's performance. So, I ended up abandoning this strategy. While I do have some target feature aggregations in my solution, as far as lags go, my final submission relies **only on lags of known features** like box score, standings, and games features.\n\n## Models used\n\nI tried out a ton of different models that didn't work. Here's a quick summary:\n\nWhat didn't work at all:\n- Denoising autoencoders\n\nWhat worked decently, but not well enough:\n- LightGBM, CatBoost\n- 1D CNN (from the [MoA competition](https://www.kaggle.com/c/lish-moa/discussion/202256))\n- Time series models on lagged targets (GRU, Transformer, DeepAR, Temporal Fusion Transformer)\n\nMy final submission ended up being an ensemble of a GRU and a Transformer that rely on non-target lags. The GRU was my best model, achieving ~1.25 on its own. The Transformer scored about ~1.27 on its own, but it proved to be very useful in an ensemble.\n\n## Validation\nI did a simple time series split. Before the training data was updated, I used July and August in 2020, and then April in 2021. After the updated data was released, I used August 2020 and June and July in 2021.\n\n## Features used\nBelow is a list of all the types of features that I used. All lagged features go back 2 weeks.\n- Mean/median/std/min/max of each target from the previous month\n- Embeddings of playerId, teamId, player position, and player status\n- Player box score lags\n    - 'gamesPlayedBatting', 'hits', 'doubles', 'triples', 'runsScored',\n    'homeRuns', 'hitByPitch', 'totalBases', 'rbi', 'stolenBases', 'assists', \n    'gamesPlayedPitching', 'completeGamesPitching', 'shutoutsPitching', 'earnedRuns', 'winsPitching', 'strikeOutsPitching', 'hitsPitching', 'saveOpportunities', 'saves', 'holds', 'inningsPitched'\n- Lagged game features\n    - 'gameTimeUTC', 'wasSigned', 'wasTraded', 'teamWins', 'teamLosses', 'teamScore', 'isHome', 'teamWon', 'scoreDiff'\n    - The game time feature only includes the hour the game was played. The idea here was a lot of the digital engagement resulting from a, say, 1 PM game would manifest itself on the day the game was played, whereas the engagement from a night game would all happen on the next day.\n- Lagged standings features\n    - 'wins', 'losses', 'pct', 'xWinLossPct', 'divisionRank', 'lastTenWins', 'lastTenLosses'\n- Lagged cumulative features (all box score features summed up by season)\n- Lagged transactional features, which were just a flag indicating if the player had either been traded or released on the particular day.\n- Player followers, team followers\n\nI tried to use features that exploited the scaling of the data, such as the figuring out the minimum increment between target features for a player and rescaling their target features to the non-scaled \"actual\" feature, but it didn't really help my model at all. I was pretty surprised by this.\n\n## Final submissions\n\nI noticed that including as much data as possible helped immensely, so I knew that I wanted to make at least one of my final solutions have models that were trained on 100% of the available data. However, my training curves for both the Transformer and GRU were [quite bumpy](https://drive.google.com/file/d/1OR2vXYhGv8W3VoDWyLDicFef79UTXjs-/view?usp=sharing). So, I took checkpoints every 200 epochs starting from step 1500 while training blind on 100% of the data. I also ensembled across several seeds.\n\nTo be safe, my second submission used models that used the last 30 days of data available as validation data. This is really just a hedge, but I'm interested to see how it does.\n\n## Other notes\n- All models were trained in PyTorch using PyTorch Lightning\n- Editor was VSCode and I made heavy use of the TensorBoard integration\n- @nyanpn's submission emulator was invaluable near the end of the competition\n- I used Optuna for hyperparameter tuning.\n- To reduce submission bugs, I used a dataset that has example versions of roster data, game data, box score data, etc. This allowed me to have a \"default\" version of the data to return when it is null.\n- As some other competitors noted, the predictions from even my best models look somewhat suspect. I was thrilled whenever I saw my model nail an upward spike in engagement. However, I think the MLB is acutely aware that some major variations in engagement happen because of exogenous factors like news stories. Given the MLB's fairly recent expansion of their social media endeavors, it makes sense that they would want to focus on purely the game-level factors that drive engagement, as this is all they really can act on. Maybe this is a faulty argument, but because of this I think MAE actually made pretty decent sense as a competition metric -- MSE would penalize you too much for factors that are completely unpredictable from game-level data.\n\nIf you made it this far, thanks for reading! I had a lot of fun during this competition and can't wait to see how things shake out.\n\nEDIT: I've made my inference notebook public here: https://www.kaggle.com/marktenenholtz/mlb-rnn-transformer-final-public\n\nA basic run-down of the models is that they're essentially the same, with the GRU and Transformer swapped in as different backbones. Other than that, they both take in the embedded categorical features as well as the lagged features, feed those into the backbone as a sequence, and then there's 3 convolutions that are applied to the output (for whatever reason, this worked better than a simple head). After the convolution, there's a set of linear layers that comprises the head of the model, and there's a unique head for each target. There are some minor differences from this in the Transformer model, but this is more or less how it works.",
      "votes": 38
    },
    {
      "id": 1423759,
      "postDate": "2021-08-03T05:55:24.283Z",
      "content": "<p><a href=\"https://www.kaggle.com/marktenenholtz\" target=\"_blank\">@marktenenholtz</a> Great work. Thanks for sharing! Do you plan to publish the Transformer code? I was thinking about implementing transformer-based model for this competition but I joined too late and didn't have time to do so. It would be great to learn from your code.</p>",
      "rawMarkdown": "@marktenenholtz Great work. Thanks for sharing! Do you plan to publish the Transformer code? I was thinking about implementing transformer-based model for this competition but I joined too late and didn't have time to do so. It would be great to learn from your code.",
      "replies": [
        {
          "id": 1436394,
          "postDate": "2021-08-03T15:06:25.503Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1436392,
          "postDate": "2021-08-03T15:06:25.503Z",
          "content": "<p>I just made my inference notebook public, and the link is in the writeup.</p>",
          "rawMarkdown": "I just made my inference notebook public, and the link is in the writeup."
        }
      ]
    },
    {
      "id": 1421584,
      "postDate": "2021-08-03T04:12:59.690Z",
      "content": "<p>Also where can I find this submission emulator you mentioned?</p>",
      "rawMarkdown": "Also where can I find this submission emulator you mentioned?",
      "replies": [
        {
          "id": 1421940,
          "postDate": "2021-08-03T04:31:58.027Z",
          "content": "<p>nvm found it <a href=\"https://www.kaggle.com/nyanpn/api-emulator-for-debugging-your-code-locally\" target=\"_blank\">https://www.kaggle.com/nyanpn/api-emulator-for-debugging-your-code-locally</a></p>",
          "rawMarkdown": "nvm found it https://www.kaggle.com/nyanpn/api-emulator-for-debugging-your-code-locally"
        }
      ]
    },
    {
      "id": 1421523,
      "postDate": "2021-08-03T04:10:04.717Z",
      "content": "<p>Thanks a lot for this post! I already learned a lot from reading this; any chance you can share your notebook? This was my first competition and I think I would learn a lot from seeing how certain things were implemented.</p>",
      "rawMarkdown": "Thanks a lot for this post! I already learned a lot from reading this; any chance you can share your notebook? This was my first competition and I think I would learn a lot from seeing how certain things were implemented.",
      "replies": [
        {
          "id": 1435929,
          "postDate": "2021-08-03T14:54:24.367Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1435921,
          "postDate": "2021-08-03T14:54:24.367Z",
          "content": "<p>Glad I could help. My code is a bit of a mess from all the craziness involved in this competition, but I'll do my best to clean it up and post it.</p>\n<p>EDIT: I decided to just make my inference notebook public. I added the link in the writeup.</p>",
          "rawMarkdown": "Glad I could help. My code is a bit of a mess from all the craziness involved in this competition, but I'll do my best to clean it up and post it.\n\nEDIT: I decided to just make my inference notebook public. I added the link in the writeup.",
          "votes": 1
        },
        {
          "id": 1449617,
          "postDate": "2021-08-04T23:26:04.527Z",
          "content": "<p>thank you!</p>",
          "rawMarkdown": "thank you!"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1423759,
      "author_name": "Zacchaeus",
      "author_url": "",
      "post_date": "2021-08-03T05:55:24.283000",
      "content": "<p><a href=\"https://www.kaggle.com/marktenenholtz\" target=\"_blank\">@marktenenholtz</a> Great work. Thanks for sharing! Do you plan to publish the Transformer code? I was thinking about implementing transformer-based model for this competition but I joined too late and didn't have time to do so. It would be great to learn from your code.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1436394,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-08-03T15:06:25.503000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1436392,
          "author_name": "Mark Tenenholtz",
          "author_url": "",
          "post_date": "2021-08-03T15:06:25.503000",
          "content": "<p>I just made my inference notebook public, and the link is in the writeup.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1421584,
      "author_name": "nateyc",
      "author_url": "",
      "post_date": "2021-08-03T04:12:59.690000",
      "content": "<p>Also where can I find this submission emulator you mentioned?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1421940,
          "author_name": "nateyc",
          "author_url": "",
          "post_date": "2021-08-03T04:31:58.027000",
          "content": "<p>nvm found it <a href=\"https://www.kaggle.com/nyanpn/api-emulator-for-debugging-your-code-locally\" target=\"_blank\">https://www.kaggle.com/nyanpn/api-emulator-for-debugging-your-code-locally</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1421523,
      "author_name": "nateyc",
      "author_url": "",
      "post_date": "2021-08-03T04:10:04.717000",
      "content": "<p>Thanks a lot for this post! I already learned a lot from reading this; any chance you can share your notebook? This was my first competition and I think I would learn a lot from seeing how certain things were implemented.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1435929,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-08-03T14:54:24.367000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1435921,
          "author_name": "Mark Tenenholtz",
          "author_url": "",
          "post_date": "2021-08-03T14:54:24.367000",
          "content": "<p>Glad I could help. My code is a bit of a mess from all the craziness involved in this competition, but I'll do my best to clean it up and post it.</p>\n<p>EDIT: I decided to just make my inference notebook public. I added the link in the writeup.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1449617,
          "author_name": "nateyc",
          "author_url": "",
          "post_date": "2021-08-04T23:26:04.527000",
          "content": "<p>thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1409239": "This was one heck of a competition. There was no shortage of new things to try at any point given the extent of the data. For that, I'd like to thank both the MLB and Kaggle.\n\n## High-level strategy\n\nIt's no secret that lagged features were very important for this competition. I found this, too, and for a while my solutions were heavily reliant on lagged target variables. I suspect most high scoring submissions were, as well. However, there was a lot of uncertainty about what the gap between the ground truth and the inference period would be. I had considered doing some augmentations for robustness, like dropping out the n most recent days of target lags randomly, but this just killed my model's performance. So, I ended up abandoning this strategy. While I do have some target feature aggregations in my solution, as far as lags go, my final submission relies **only on lags of known features** like box score, standings, and games features.\n\n## Models used\n\nI tried out a ton of different models that didn't work. Here's a quick summary:\n\nWhat didn't work at all:\n- Denoising autoencoders\n\nWhat worked decently, but not well enough:\n- LightGBM, CatBoost\n- 1D CNN (from the [MoA competition](https://www.kaggle.com/c/lish-moa/discussion/202256))\n- Time series models on lagged targets (GRU, Transformer, DeepAR, Temporal Fusion Transformer)\n\nMy final submission ended up being an ensemble of a GRU and a Transformer that rely on non-target lags. The GRU was my best model, achieving ~1.25 on its own. The Transformer scored about ~1.27 on its own, but it proved to be very useful in an ensemble.\n\n## Validation\nI did a simple time series split. Before the training data was updated, I used July and August in 2020, and then April in 2021. After the updated data was released, I used August 2020 and June and July in 2021.\n\n## Features used\nBelow is a list of all the types of features that I used. All lagged features go back 2 weeks.\n- Mean/median/std/min/max of each target from the previous month\n- Embeddings of playerId, teamId, player position, and player status\n- Player box score lags\n    - 'gamesPlayedBatting', 'hits', 'doubles', 'triples', 'runsScored',\n    'homeRuns', 'hitByPitch', 'totalBases', 'rbi', 'stolenBases', 'assists', \n    'gamesPlayedPitching', 'completeGamesPitching', 'shutoutsPitching', 'earnedRuns', 'winsPitching', 'strikeOutsPitching', 'hitsPitching', 'saveOpportunities', 'saves', 'holds', 'inningsPitched'\n- Lagged game features\n    - 'gameTimeUTC', 'wasSigned', 'wasTraded', 'teamWins', 'teamLosses', 'teamScore', 'isHome', 'teamWon', 'scoreDiff'\n    - The game time feature only includes the hour the game was played. The idea here was a lot of the digital engagement resulting from a, say, 1 PM game would manifest itself on the day the game was played, whereas the engagement from a night game would all happen on the next day.\n- Lagged standings features\n    - 'wins', 'losses', 'pct', 'xWinLossPct', 'divisionRank', 'lastTenWins', 'lastTenLosses'\n- Lagged cumulative features (all box score features summed up by season)\n- Lagged transactional features, which were just a flag indicating if the player had either been traded or released on the particular day.\n- Player followers, team followers\n\nI tried to use features that exploited the scaling of the data, such as the figuring out the minimum increment between target features for a player and rescaling their target features to the non-scaled \"actual\" feature, but it didn't really help my model at all. I was pretty surprised by this.\n\n## Final submissions\n\nI noticed that including as much data as possible helped immensely, so I knew that I wanted to make at least one of my final solutions have models that were trained on 100% of the available data. However, my training curves for both the Transformer and GRU were [quite bumpy](https://drive.google.com/file/d/1OR2vXYhGv8W3VoDWyLDicFef79UTXjs-/view?usp=sharing). So, I took checkpoints every 200 epochs starting from step 1500 while training blind on 100% of the data. I also ensembled across several seeds.\n\nTo be safe, my second submission used models that used the last 30 days of data available as validation data. This is really just a hedge, but I'm interested to see how it does.\n\n## Other notes\n- All models were trained in PyTorch using PyTorch Lightning\n- Editor was VSCode and I made heavy use of the TensorBoard integration\n- @nyanpn's submission emulator was invaluable near the end of the competition\n- I used Optuna for hyperparameter tuning.\n- To reduce submission bugs, I used a dataset that has example versions of roster data, game data, box score data, etc. This allowed me to have a \"default\" version of the data to return when it is null.\n- As some other competitors noted, the predictions from even my best models look somewhat suspect. I was thrilled whenever I saw my model nail an upward spike in engagement. However, I think the MLB is acutely aware that some major variations in engagement happen because of exogenous factors like news stories. Given the MLB's fairly recent expansion of their social media endeavors, it makes sense that they would want to focus on purely the game-level factors that drive engagement, as this is all they really can act on. Maybe this is a faulty argument, but because of this I think MAE actually made pretty decent sense as a competition metric -- MSE would penalize you too much for factors that are completely unpredictable from game-level data.\n\nIf you made it this far, thanks for reading! I had a lot of fun during this competition and can't wait to see how things shake out.\n\nEDIT: I've made my inference notebook public here: https://www.kaggle.com/marktenenholtz/mlb-rnn-transformer-final-public\n\nA basic run-down of the models is that they're essentially the same, with the GRU and Transformer swapped in as different backbones. Other than that, they both take in the embedded categorical features as well as the lagged features, feed those into the backbone as a sequence, and then there's 3 convolutions that are applied to the output (for whatever reason, this worked better than a simple head). After the convolution, there's a set of linear layers that comprises the head of the model, and there's a unique head for each target. There are some minor differences from this in the Transformer model, but this is more or less how it works.",
    "1423759": "@marktenenholtz Great work. Thanks for sharing! Do you plan to publish the Transformer code? I was thinking about implementing transformer-based model for this competition but I joined too late and didn't have time to do so. It would be great to learn from your code.",
    "1421584": "Also where can I find this submission emulator you mentioned?",
    "1421523": "Thanks a lot for this post! I already learned a lot from reading this; any chance you can share your notebook? This was my first competition and I think I would learn a lot from seeing how certain things were implemented."
  }
}