{
  "id": 256620,
  "title": "[3rd place] Summary of my work",
  "url": "/competitions/mlb-player-digital-engagement-forecasting/writeups/nyanp-w-o-leak-1-3004-3rd-place-summary-of-my-work",
  "author_name": "",
  "post_date": "2021-09-09T22:56:23.847Z",
  "votes": 91,
  "comment_count": 5,
  "views": 0,
  "content": "<p><strong><em>2021-09-10 updated</em></strong><br>\nThe final ranking is 3rd! Thank you all :) My source code is a bit messy, but you can see it here: <a href=\"https://github.com/nyanp/mlb-player-digital-engagement\" target=\"_blank\">https://github.com/nyanp/mlb-player-digital-engagement</a> and <a href=\"https://www.kaggle.com/nyanpn/3rd-place-solution-inference-only\" target=\"_blank\">https://www.kaggle.com/nyanpn/3rd-place-solution-inference-only</a> .</p>\n<hr>\n<p>First of all, I would like to thank the hosts for organizing this competition.<br>\nIt was a very tough competition but I learned a lot from it.</p>\n<p>I have no idea what my final standings will be, but I will share with you what I did.</p>\n<h2>Model strategy</h2>\n<p>Lag features are very useful in this competition, but we cannot use target information in the test data period. So we need to set an appropriate \"gap\" between the prediction date and the lag feature data.</p>\n<p>For example, if you want to make a prediction for 8/4, <br>\nyou need to create a gap of at least 3 days because you cannot use the target information from 8/1 to 8/3.</p>\n<p>I divided the test data into 8 periods and used different models with different gaps in each period (gap: 0, 3, 7, 14, 21, 28, 35 ,45days).</p>\n<p><img src=\"https://raw.githubusercontent.com/nyanp/mlb-player-digital-engagement/main/docs/img/time%20series%20gap.png\" alt=\"https://raw.githubusercontent.com/nyanp/mlb-player-digital-engagement/main/docs/img/time%20series%20gap.png\"></p>\n<h2>Model</h2>\n<p>The three models were ensembled with different weights for each target.</p>\n<ul>\n<li>LightGBM</li>\n<li>MLP</li>\n<li>1DCNN</li>\n</ul>\n<p>1DCNN is the same as <a href=\"https://www.kaggle.com/c/lish-moa/discussion/202256\" target=\"_blank\">the 2nd place solution of MoA Competition</a>.</p>\n<h3>Validation</h3>\n<p>Time based split.</p>\n<ul>\n<li>validation before update: 2019/8, 2020/8, 2021/4</li>\n<li>validation after update: 2020/8, 2021/6, 2021/7</li>\n</ul>\n<p>Basically, I only adopted ideas that improved the score in all periods.</p>\n<h2>Features</h2>\n<p>I used ~440 features. In addition to joins and asof merge of basic tables, the following features were used:</p>\n<ul>\n<li>lag features per player<ul>\n<li>Average of the last 7/28/70/360/720 days</li>\n<li>Average over the on-seasons</li>\n<li>Average for the same period in the previous year</li>\n<li>Average of days with/without a game</li></ul></li>\n<li>number of events, pitch events, action events</li>\n<li>days from last rosters, awards, transactions and box scores</li>\n<li>sum of box scores in the last 7/30/90 days</li>\n<li>number of games and events in the day</li>\n<li>event-level meta feature<ul>\n<li>aggregation of predictions of model trained on event table</li>\n<li>group by (date, playerId), (date, teamId) and (date)</li></ul></li>\n</ul>\n<h3>Cumcount Leakage</h3>\n<p>There is a strange correlation between the cumcount of the dataframe retrieved from the Time-Series API and the target.</p>\n<p>I noticed this problem 3 days before the competition ended. I did not post it in the discussion as it might confuse the participants, but contacted the host immediately.<br>\nAdding this cumcount to the features only improves the CV a little bit, so it's probably some kind of artifact or something, but even if it doesn't improve the CV much, it's better to shuffle the test data since it's nonsense that the order of the rows makes sense.</p>\n<p>I did not end up using this leak for final submission.</p>\n<h3>Implementation Note</h3>\n<p>Building a complex data pipeline in Jupyter Notebook with the Time Series API can be a big pain. I'll share some of my efforts.</p>\n<ul>\n<li>Maintain the source code on GitHub and paste the BASE64-encoded code into the jupyter notebook<ul>\n<li>see: <a href=\"https://github.com/lopuhin/kaggle-imet-2019\" target=\"_blank\">https://github.com/lopuhin/kaggle-imet-2019</a></li></ul></li>\n<li>The inference notebook is also maintained on GitHub and automatically uploaded as the Kaggle Kernel through GitHub Actions</li>\n<li>Avoid the use of pandas and instead use a dictionary of numpy arrays to manage state updates</li>\n<li>Use the same feature generation function for training data and inference<ul>\n<li>Both training and test are treated as streaming data, and features were generated using for-loop.</li>\n<li>This is the most important point to get a stable and bug-free data pipeline</li>\n<li>see: <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/196942\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/196942</a></li></ul></li>\n<li>Debug code locally using the API emulator<ul>\n<li>Test the robustness of my inference pipeline by \"dropout\" some of the data returned by the emulator (a kind of \"Chaos Engineering\")</li></ul></li>\n<li>Catch exceptions in various functions and convert them to appropriate \"default\" values</li>\n</ul>\n<p>Thanks to all of this, I was able to finish the competition 1st stage without making a single submission error.</p>",
  "messages": [
    {
      "id": "1408712",
      "postDate": "08/02/2021 15:16:17",
      "content": "<p><strong><em>2021-09-10 updated</em></strong><br>\nThe final ranking is 3rd! Thank you all :) My source code is a bit messy, but you can see it here: <a href=\"https://github.com/nyanp/mlb-player-digital-engagement\" target=\"_blank\">https://github.com/nyanp/mlb-player-digital-engagement</a> and <a href=\"https://www.kaggle.com/nyanpn/3rd-place-solution-inference-only\" target=\"_blank\">https://www.kaggle.com/nyanpn/3rd-place-solution-inference-only</a> .</p>\n<hr>\n<p>First of all, I would like to thank the hosts for organizing this competition.<br>\nIt was a very tough competition but I learned a lot from it.</p>\n<p>I have no idea what my final standings will be, but I will share with you what I did.</p>\n<h2>Model strategy</h2>\n<p>Lag features are very useful in this competition, but we cannot use target information in the test data period. So we need to set an appropriate \"gap\" between the prediction date and the lag feature data.</p>\n<p>For example, if you want to make a prediction for 8/4, <br>\nyou need to create a gap of at least 3 days because you cannot use the target information from 8/1 to 8/3.</p>\n<p>I divided the test data into 8 periods and used different models with different gaps in each period (gap: 0, 3, 7, 14, 21, 28, 35 ,45days).</p>\n<p><img src=\"https://raw.githubusercontent.com/nyanp/mlb-player-digital-engagement/main/docs/img/time%20series%20gap.png\" alt=\"https://raw.githubusercontent.com/nyanp/mlb-player-digital-engagement/main/docs/img/time%20series%20gap.png\"></p>\n<h2>Model</h2>\n<p>The three models were ensembled with different weights for each target.</p>\n<ul>\n<li>LightGBM</li>\n<li>MLP</li>\n<li>1DCNN</li>\n</ul>\n<p>1DCNN is the same as <a href=\"https://www.kaggle.com/c/lish-moa/discussion/202256\" target=\"_blank\">the 2nd place solution of MoA Competition</a>.</p>\n<h3>Validation</h3>\n<p>Time based split.</p>\n<ul>\n<li>validation before update: 2019/8, 2020/8, 2021/4</li>\n<li>validation after update: 2020/8, 2021/6, 2021/7</li>\n</ul>\n<p>Basically, I only adopted ideas that improved the score in all periods.</p>\n<h2>Features</h2>\n<p>I used ~440 features. In addition to joins and asof merge of basic tables, the following features were used:</p>\n<ul>\n<li>lag features per player<ul>\n<li>Average of the last 7/28/70/360/720 days</li>\n<li>Average over the on-seasons</li>\n<li>Average for the same period in the previous year</li>\n<li>Average of days with/without a game</li></ul></li>\n<li>number of events, pitch events, action events</li>\n<li>days from last rosters, awards, transactions and box scores</li>\n<li>sum of box scores in the last 7/30/90 days</li>\n<li>number of games and events in the day</li>\n<li>event-level meta feature<ul>\n<li>aggregation of predictions of model trained on event table</li>\n<li>group by (date, playerId), (date, teamId) and (date)</li></ul></li>\n</ul>\n<h3>Cumcount Leakage</h3>\n<p>There is a strange correlation between the cumcount of the dataframe retrieved from the Time-Series API and the target.</p>\n<p>I noticed this problem 3 days before the competition ended. I did not post it in the discussion as it might confuse the participants, but contacted the host immediately.<br>\nAdding this cumcount to the features only improves the CV a little bit, so it's probably some kind of artifact or something, but even if it doesn't improve the CV much, it's better to shuffle the test data since it's nonsense that the order of the rows makes sense.</p>\n<p>I did not end up using this leak for final submission.</p>\n<h3>Implementation Note</h3>\n<p>Building a complex data pipeline in Jupyter Notebook with the Time Series API can be a big pain. I'll share some of my efforts.</p>\n<ul>\n<li>Maintain the source code on GitHub and paste the BASE64-encoded code into the jupyter notebook<ul>\n<li>see: <a href=\"https://github.com/lopuhin/kaggle-imet-2019\" target=\"_blank\">https://github.com/lopuhin/kaggle-imet-2019</a></li></ul></li>\n<li>The inference notebook is also maintained on GitHub and automatically uploaded as the Kaggle Kernel through GitHub Actions</li>\n<li>Avoid the use of pandas and instead use a dictionary of numpy arrays to manage state updates</li>\n<li>Use the same feature generation function for training data and inference<ul>\n<li>Both training and test are treated as streaming data, and features were generated using for-loop.</li>\n<li>This is the most important point to get a stable and bug-free data pipeline</li>\n<li>see: <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/196942\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/196942</a></li></ul></li>\n<li>Debug code locally using the API emulator<ul>\n<li>Test the robustness of my inference pipeline by \"dropout\" some of the data returned by the emulator (a kind of \"Chaos Engineering\")</li></ul></li>\n<li>Catch exceptions in various functions and convert them to appropriate \"default\" values</li>\n</ul>\n<p>Thanks to all of this, I was able to finish the competition 1st stage without making a single submission error.</p>",
      "rawMarkdown": "***2021-09-10 updated***\nThe final ranking is 3rd! Thank you all :) My source code is a bit messy, but you can see it here: https://github.com/nyanp/mlb-player-digital-engagement and https://www.kaggle.com/nyanpn/3rd-place-solution-inference-only .\n\n---\n\nFirst of all, I would like to thank the hosts for organizing this competition.\nIt was a very tough competition but I learned a lot from it.\n\nI have no idea what my final standings will be, but I will share with you what I did.\n\n## Model strategy\nLag features are very useful in this competition, but we cannot use target information in the test data period. So we need to set an appropriate \"gap\" between the prediction date and the lag feature data.\n\nFor example, if you want to make a prediction for 8/4, \nyou need to create a gap of at least 3 days because you cannot use the target information from 8/1 to 8/3.\n\nI divided the test data into 8 periods and used different models with different gaps in each period (gap: 0, 3, 7, 14, 21, 28, 35 ,45days).\n\n![https://raw.githubusercontent.com/nyanp/mlb-player-digital-engagement/main/docs/img/time%20series%20gap.png](https://raw.githubusercontent.com/nyanp/mlb-player-digital-engagement/main/docs/img/time%20series%20gap.png)\n\n## Model\nThe three models were ensembled with different weights for each target.\n\n- LightGBM\n- MLP\n- 1DCNN\n\n1DCNN is the same as [the 2nd place solution of MoA Competition](https://www.kaggle.com/c/lish-moa/discussion/202256).\n\n### Validation\nTime based split.\n- validation before update: 2019/8, 2020/8, 2021/4\n- validation after update: 2020/8, 2021/6, 2021/7\n\nBasically, I only adopted ideas that improved the score in all periods.\n\n## Features\nI used ~440 features. In addition to joins and asof merge of basic tables, the following features were used:\n\n- lag features per player\n    - Average of the last 7/28/70/360/720 days\n    - Average over the on-seasons\n    - Average for the same period in the previous year\n    - Average of days with/without a game\n- number of events, pitch events, action events\n- days from last rosters, awards, transactions and box scores\n- sum of box scores in the last 7/30/90 days\n- number of games and events in the day\n- event-level meta feature\n    - aggregation of predictions of model trained on event table\n    - group by (date, playerId), (date, teamId) and (date)\n\n### Cumcount Leakage\nThere is a strange correlation between the cumcount of the dataframe retrieved from the Time-Series API and the target.\n\nI noticed this problem 3 days before the competition ended. I did not post it in the discussion as it might confuse the participants, but contacted the host immediately.\nAdding this cumcount to the features only improves the CV a little bit, so it's probably some kind of artifact or something, but even if it doesn't improve the CV much, it's better to shuffle the test data since it's nonsense that the order of the rows makes sense.\n\nI did not end up using this leak for final submission.\n\n### Implementation Note\nBuilding a complex data pipeline in Jupyter Notebook with the Time Series API can be a big pain. I'll share some of my efforts.\n\n- Maintain the source code on GitHub and paste the BASE64-encoded code into the jupyter notebook\n    - see: https://github.com/lopuhin/kaggle-imet-2019\n- The inference notebook is also maintained on GitHub and automatically uploaded as the Kaggle Kernel through GitHub Actions\n- Avoid the use of pandas and instead use a dictionary of numpy arrays to manage state updates\n- Use the same feature generation function for training data and inference\n    - Both training and test are treated as streaming data, and features were generated using for-loop.\n    - This is the most important point to get a stable and bug-free data pipeline\n    - see: https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/196942\n- Debug code locally using the API emulator\n    - Test the robustness of my inference pipeline by \"dropout\" some of the data returned by the emulator (a kind of \"Chaos Engineering\")\n- Catch exceptions in various functions and convert them to appropriate \"default\" values\n\nThanks to all of this, I was able to finish the competition 1st stage without making a single submission error.",
      "votes": null
    },
    {
      "id": "1408746",
      "postDate": "08/02/2021 15:28:05",
      "content": "<p>I’ve shared the notebook about cumcount leakage.<br>\n<a href=\"https://www.kaggle.com/nyanpn/cumcount-leakage\" target=\"_blank\">https://www.kaggle.com/nyanpn/cumcount-leakage</a></p>",
      "rawMarkdown": "I’ve shared the notebook about cumcount leakage.\nhttps://www.kaggle.com/nyanpn/cumcount-leakage",
      "votes": null
    },
    {
      "id": "1408755",
      "postDate": "08/02/2021 15:32:55",
      "content": "<p>I also found that the 1D CNN that you mentioned worked pretty well for me, although it didn't make it into my final submission. Thanks for the writeup!</p>",
      "rawMarkdown": "I also found that the 1D CNN that you mentioned worked pretty well for me, although it didn't make it into my final submission. Thanks for the writeup!",
      "votes": null
    },
    {
      "id": "1502702",
      "postDate": "09/04/2021 14:53:40",
      "content": "<p>Hello, how do you deal with nan problem with this 1D CNN?</p>",
      "rawMarkdown": "Hello, how do you deal with nan problem with this 1D CNN?",
      "votes": null
    },
    {
      "id": "1508189",
      "postDate": "09/10/2021 01:41:12",
      "content": "<p>Awesome, great job <a href=\"https://www.kaggle.com/nyanpn\" target=\"_blank\">@nyanpn</a> . Congrats on solo gold 3rd place!</p>",
      "rawMarkdown": "Awesome, great job @nyanpn . Congrats on solo gold 3rd place!",
      "votes": null
    },
    {
      "id": "1508741",
      "postDate": "09/10/2021 14:16:41",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/nyanpn\" target=\"_blank\">@nyanpn</a> . Use of github actions to streamline inference pipeline is really creative.👍</p>",
      "rawMarkdown": "Congratulations @nyanpn . Use of github actions to streamline inference pipeline is really creative.👍",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1408746,
      "author_name": "nyanpn",
      "author_url": "",
      "post_date": "08/02/2021 15:28:05",
      "content": "<p>I’ve shared the notebook about cumcount leakage.<br>\n<a href=\"https://www.kaggle.com/nyanpn/cumcount-leakage\" target=\"_blank\">https://www.kaggle.com/nyanpn/cumcount-leakage</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1408755,
      "author_name": "marktenenholtz",
      "author_url": "",
      "post_date": "08/02/2021 15:32:55",
      "content": "<p>I also found that the 1D CNN that you mentioned worked pretty well for me, although it didn't make it into my final submission. Thanks for the writeup!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1502702,
          "author_name": "yuanzhezhou",
          "author_url": "",
          "post_date": "09/04/2021 14:53:40",
          "content": "<p>Hello, how do you deal with nan problem with this 1D CNN?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1508189,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "09/10/2021 01:41:12",
      "content": "<p>Awesome, great job <a href=\"https://www.kaggle.com/nyanpn\" target=\"_blank\">@nyanpn</a> . Congrats on solo gold 3rd place!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1508741,
      "author_name": "tezdhar",
      "author_url": "",
      "post_date": "09/10/2021 14:16:41",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/nyanpn\" target=\"_blank\">@nyanpn</a> . Use of github actions to streamline inference pipeline is really creative.👍</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1408712": "***2021-09-10 updated***\nThe final ranking is 3rd! Thank you all :) My source code is a bit messy, but you can see it here: https://github.com/nyanp/mlb-player-digital-engagement and https://www.kaggle.com/nyanpn/3rd-place-solution-inference-only .\n\n---\n\nFirst of all, I would like to thank the hosts for organizing this competition.\nIt was a very tough competition but I learned a lot from it.\n\nI have no idea what my final standings will be, but I will share with you what I did.\n\n## Model strategy\nLag features are very useful in this competition, but we cannot use target information in the test data period. So we need to set an appropriate \"gap\" between the prediction date and the lag feature data.\n\nFor example, if you want to make a prediction for 8/4, \nyou need to create a gap of at least 3 days because you cannot use the target information from 8/1 to 8/3.\n\nI divided the test data into 8 periods and used different models with different gaps in each period (gap: 0, 3, 7, 14, 21, 28, 35 ,45days).\n\n![https://raw.githubusercontent.com/nyanp/mlb-player-digital-engagement/main/docs/img/time%20series%20gap.png](https://raw.githubusercontent.com/nyanp/mlb-player-digital-engagement/main/docs/img/time%20series%20gap.png)\n\n## Model\nThe three models were ensembled with different weights for each target.\n\n- LightGBM\n- MLP\n- 1DCNN\n\n1DCNN is the same as [the 2nd place solution of MoA Competition](https://www.kaggle.com/c/lish-moa/discussion/202256).\n\n### Validation\nTime based split.\n- validation before update: 2019/8, 2020/8, 2021/4\n- validation after update: 2020/8, 2021/6, 2021/7\n\nBasically, I only adopted ideas that improved the score in all periods.\n\n## Features\nI used ~440 features. In addition to joins and asof merge of basic tables, the following features were used:\n\n- lag features per player\n    - Average of the last 7/28/70/360/720 days\n    - Average over the on-seasons\n    - Average for the same period in the previous year\n    - Average of days with/without a game\n- number of events, pitch events, action events\n- days from last rosters, awards, transactions and box scores\n- sum of box scores in the last 7/30/90 days\n- number of games and events in the day\n- event-level meta feature\n    - aggregation of predictions of model trained on event table\n    - group by (date, playerId), (date, teamId) and (date)\n\n### Cumcount Leakage\nThere is a strange correlation between the cumcount of the dataframe retrieved from the Time-Series API and the target.\n\nI noticed this problem 3 days before the competition ended. I did not post it in the discussion as it might confuse the participants, but contacted the host immediately.\nAdding this cumcount to the features only improves the CV a little bit, so it's probably some kind of artifact or something, but even if it doesn't improve the CV much, it's better to shuffle the test data since it's nonsense that the order of the rows makes sense.\n\nI did not end up using this leak for final submission.\n\n### Implementation Note\nBuilding a complex data pipeline in Jupyter Notebook with the Time Series API can be a big pain. I'll share some of my efforts.\n\n- Maintain the source code on GitHub and paste the BASE64-encoded code into the jupyter notebook\n    - see: https://github.com/lopuhin/kaggle-imet-2019\n- The inference notebook is also maintained on GitHub and automatically uploaded as the Kaggle Kernel through GitHub Actions\n- Avoid the use of pandas and instead use a dictionary of numpy arrays to manage state updates\n- Use the same feature generation function for training data and inference\n    - Both training and test are treated as streaming data, and features were generated using for-loop.\n    - This is the most important point to get a stable and bug-free data pipeline\n    - see: https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/196942\n- Debug code locally using the API emulator\n    - Test the robustness of my inference pipeline by \"dropout\" some of the data returned by the emulator (a kind of \"Chaos Engineering\")\n- Catch exceptions in various functions and convert them to appropriate \"default\" values\n\nThanks to all of this, I was able to finish the competition 1st stage without making a single submission error.",
    "1408746": "I’ve shared the notebook about cumcount leakage.\nhttps://www.kaggle.com/nyanpn/cumcount-leakage",
    "1408755": "I also found that the 1D CNN that you mentioned worked pretty well for me, although it didn't make it into my final submission. Thanks for the writeup!",
    "1502702": "Hello, how do you deal with nan problem with this 1D CNN?",
    "1508189": "Awesome, great job @nyanpn . Congrats on solo gold 3rd place!",
    "1508741": "Congratulations @nyanpn . Use of github actions to streamline inference pipeline is really creative.👍"
  },
  "source": "meta"
}