{
  "id": 271890,
  "title": "6th solution",
  "url": "/competitions/mlb-player-digital-engagement-forecasting/writeups/npb-6th-solution",
  "author_name": "",
  "post_date": "2021-09-13T03:22:50.668404800Z",
  "votes": 19,
  "comment_count": 3,
  "views": 0,
  "content": "<p>First of all, I also would like to thank this competition host and all competitors. It has been about two years since I got my first silver medal in the APTOS competition, and I'm finally a Kaggle master. I'm very pleased now. I'm still new, but I'll continue to learn from Kaggle with a sincere attitude.<br>\nI am also happy that I could get a gold medal in a competition related to my favorite sport what is baseball. Baseball may be a minor sport on the global stage, but it's very popular in Japan, and I watch NPB (Nippon Professional Baseball Organization) games every week. I wish this fascinating and strategic would become more popular around the world!</p>\n<h1>To member (<a href=\"https://www.kaggle.com/tea1013\" target=\"_blank\">tea</a> &amp; <a href=\"https://www.kaggle.com/nomorevotch\" target=\"_blank\">sqrt4kaido</a>)</h1>\n<p>I've been participated in competitions as a solo, and this MLB competition was my first experience participating as a team. I learned a lot of things when I participated competitions by solo, but I learned even more this time. I could get a gold medal thank to your a lot of ideas.</p>\n<h1>tea's part</h1>\n<h2>Models</h2>\n<ul>\n<li>LGBM  lag / no lag</li>\n<li>CatBoost  lag / no lag<br>\nUse <strong>optuna</strong> to optimize hyper-parameters.</li>\n</ul>\n<h2>Train &amp; Valid</h2>\n<ul>\n<li>Use 2018-01-01 ~ 2021-05-31 data for train, and 2021-06-01 ~ 2021-07-17 for validation</li>\n<li>During the test period, I use 2018-01-01 ~ 2021-06-30 data for train</li>\n<li>Limit data to in-season data only</li>\n</ul>\n<h2>Features mainly effective</h2>\n<ul>\n<li>target lags (for 45 days)</li>\n<li>player target statistics (mean, median, max, min, var, skew, kurt)<ul>\n<li>Use the respective statistics for April, May, and June 2021.</li>\n<li>Also use the respective statistics for game day and no game day.</li></ul></li>\n<li>team target statistics (mean, median, max, min, var, skew, kurt)<ul>\n<li>Use statistics for June 2021 only</li></ul></li>\n<li>daysSinceLastGame / Roster</li>\n<li>days from the beginning of the year / month</li>\n<li>day of week</li>\n<li>years from the debut year</li>\n<li>age</li>\n<li>position</li>\n<li>player status</li>\n<li>playerBoxScores features<br>\n<br></li>\n</ul>\n<h1>sqrt4kaido part</h1>\n<h2>Models</h2>\n<ul>\n<li>LGBM only no lag</li>\n</ul>\n<ol>\n<li>The first seed, only the players in test are used for training.</li>\n<li>Second seed, only the players in test are used for training.</li>\n<li>The first seed. All the players are used for training.<br>\nUse <strong>optuna</strong> to optimize hyper-parameters.<br>\nBefore update train.csv, the best score for my single model was 1.3146.</li>\n</ol>\n<h2>Train &amp; Valid</h2>\n<ul>\n<li>training phase<br>\nUse 2018-01-01 ~ 2021-03-31 data for train, and 2021-04-01 ~ 2021-04-30 for validation</li>\n<li>after update train.csv<br>\nUse 2018-01-01 ~ 2021-05-31 data for train, and 2021-06-01 ~ 2021-06-30 for validation<br>\nI did not use the data for the first half of July because, unlike the other months, the values were very small.</li>\n<li>Limit data to in-season data only &lt;- important</li>\n</ul>\n<h2>Features mainly effective</h2>\n<ul>\n<li>player target statistics (mean, median, max, min, var, skew, kurt)<br>\nUse the respective statistics for June 2021. This feature is leaked to the valid data, but we used it because it was also valid for public testing.</li>\n<li>team target statistics (mean, median, max, min, var, skew, kurt)<br>\nsame as above</li>\n<li>season info<br>\npre/post season? regular season? alltardate? etc.</li>\n<li>award flag</li>\n<li>daysSinceLastGame / Roster &lt;- important</li>\n<li>day of week</li>\n<li>years from the debut year</li>\n<li>age</li>\n<li>position</li>\n<li>player status</li>\n<li>playerBoxScores features</li>\n</ul>\n<h2>did not work</h2>\n<ul>\n<li>Create separate models for the data with and without matches.<br>\nAfter update train.csv, it stopped working for some reason.</li>\n<li>Undo the scale and learn<br>\nThis time the target is scaled from 0 to 100, but there was a <a href=\"https://www.kaggle.com/c/mlb-player-digital-engagement-forecasting/discussion/253940\" target=\"_blank\">discussion</a> that it seems to be max-scaled for each day.<br>\nSo I tried training with the scale back, and the CV results improved. However, when we actually applied it to the test data, the score worsened because the wrong scaling was applied when the max prediction was shifted, so all the predictions for the day were shifted.<br>\nI also tried scaling to the average of the same month in the previous year, but this did not improve the accuracy.<br>\n<br></li>\n</ul>\n<h1>Makabe's part</h1>\n<h2>Models</h2>\n<ul>\n<li>I used three models.<ul>\n<li>Simple NN (add lags)</li>\n<li>Light GBM (add lags)</li>\n<li>Light GBM (none lags)</li></ul></li>\n<li>The reason to use three type models is the following.<ul>\n<li>To ensemble is useful making robust models.</li>\n<li>Regarding the lag feature, <strong>which is the key to this competition</strong>, I noticed the following fact in the middle of the competition.</li>\n<li>(If we can get the correct target information as a lag early in the test period.)<ul>\n<li>Light GBM (add lags) &gt; Simple NN (add lags) &gt; Light GBM (none lags)</li></ul></li>\n<li>(If we have to use the target information predicted by model as lag late in the test period.)<ul>\n<li>Light GBM (none lags) &gt; Simple NN (add lags) &gt; Light GBM (add lags)</li></ul></li>\n<li>From the above, I can see that lag feature is very effective if it contains the correct values. However, because of its effectiveness, there is also a possibility of overfitting. This tendency was especially for LGB.</li>\n<li>I decided to control the blend ratio of the three models according to the number of days in test period to balance lag features (effectiveness &amp; overfitting).</li>\n<li>The method of controlling the size of lag features depending on test time period was pointed out by <a href=\"https://www.kaggle.com/c/mlb-player-digital-engagement-forecasting/discussion/256620\" target=\"_blank\">other Kagglers</a> too. It was very effective method and we could this knowledge for our entire team ensemble too.</li></ul></li>\n</ul>\n<h2>Training / Validation</h2>\n<ul>\n<li>I used 5 folds.</li>\n</ul>\n<ol>\n<li>&lt; 20200801(training) &amp; 2020/08/01 - 2020/09/30(validation)</li>\n<li>&lt; 20200901(training) &amp; 2020/09/01 - 2020/10/31(validation)</li>\n<li>&lt; 20210401(training) &amp; 2021/04/01 - 2021/05/31(validation)</li>\n<li>&lt; 20210501(training) &amp; 2021/05/01 - 2021/06/30(validation)</li>\n<li>&lt; 20210601(training) &amp; 2021/06/01 - 2021/07/xx(validation)</li>\n</ol>\n<ul>\n<li>I decided on period with the following two points in mind.<ul>\n<li>Season with irregular schedule due to Covid-19 (2020).</li>\n<li>The trend of target feature is very different between the first and second half of the season.</li></ul></li>\n<li>However, I hadn't been able to determine the period above based on a deep consideration of this impact. I think there is room for reconsideration.</li>\n</ul>\n<h2>Features</h2>\n<p>The following is a list of feature what is helped improve accuracy. (There were many useful features than the below, but most of them were already in other Kaggler's notebooks, so I omit them.)</p>\n<ul>\n<li>stats of target feature<ul>\n<li>I labeled the data by period as follows, and I added stats that is the one before as features.</li>\n<li>2018.3 &amp; 2018.4 (=label1), 2018.5 &amp; 2018.6 (=label2), …</li>\n<li>I thought that by doing this, and I could use stats of target until September 2021. However, I knew after the competition, test period was until August 31, so I didn't need to use the stats of target two months earlier.</li>\n<li>I used mean, median, distribution,.. and more.</li>\n<li>I calculated the stats of targets by defense position, roster status, and team, respectively.</li>\n<li>There was difference of targets by roster status. (The difference was especially noticeable in target1.)</li></ul></li>\n<li>number of days what have elapsed since the last time player participated in game.<ul>\n<li>This ideas was from <a href=\"https://www.kaggle.com/nomorevotch\" target=\"_blank\">@sqrt4kaido</a>. This variable was based on the fact that the engagement of players who haven't played for a long time tends to decline over time, and contributed to improving accuracy.</li></ul></li>\n<li>score from the day before yesterday<ul>\n<li>I thought the success of a game is continuous, and player's popularity is not necessarily based only on the previous day's performance, but past success is a factor.</li>\n<li><a href=\"https://www.kaggle.com/c/mlb-player-digital-engagement-forecasting/discussion/266596\" target=\"_blank\">Team AutoMLB seems to this point</a>, and incorporated player scores into the model for a considerable period of time. I probably should have included more time periods than just the day before yesterday.</li></ul></li>\n<li>A lot of stats<ul>\n<li>I used various features based on <a href=\"https://sabr.org/sabermetrics\" target=\"_blank\">Sabermetric</a>. (K/BB, RC, OPS, FIP.. and more)</li>\n<li>This is especially for pitchers, I thought the ratio divided by the number of innings pitched or the number of pitches thrown tends to be a better representation of the player than the absolute number. (It's no surprise that starting pitchers strike out more than relief pitchers.) For this reason, I also incorporate many indicators divided by the number of innings or throw of pitches(K9, HR9.. and more).<br>\n<br></li></ul></li>\n</ul>\n<h1>Doing by team</h1>\n<ul>\n<li>In order to ensemble easily, it was necessary to improve the reproducibility and reusability of each member’s source code. Therefore, we unified the specifications and developed the prediction classes with common methods.</li>\n<li>In this competition, it was very important to finsh the process correctly, because the Time Series API is very complicated. Therefore, we created a LocalTestClass to reproduce the Time Series API in local environment.</li>\n<li>By using this class, we could verify the accuracy in a larger period.</li>\n<li>We controlled the ensemble ratio for the fact that the effectiveness of lag feature weakens over time.<br>\n<a href=\"https://postimg.cc/gwwHCDJm\" target=\"_blank\"><img src=\"https://i.postimg.cc/g0st6tbX/MLB-sub.png\" alt=\"MLB-sub.png\"></a></li>\n</ul>",
  "messages": [
    {
      "id": "1511022",
      "postDate": "09/13/2021 03:22:50",
      "content": "<p>First of all, I also would like to thank this competition host and all competitors. It has been about two years since I got my first silver medal in the APTOS competition, and I'm finally a Kaggle master. I'm very pleased now. I'm still new, but I'll continue to learn from Kaggle with a sincere attitude.<br>\nI am also happy that I could get a gold medal in a competition related to my favorite sport what is baseball. Baseball may be a minor sport on the global stage, but it's very popular in Japan, and I watch NPB (Nippon Professional Baseball Organization) games every week. I wish this fascinating and strategic would become more popular around the world!</p>\n<h1>To member (<a href=\"https://www.kaggle.com/tea1013\" target=\"_blank\">tea</a> &amp; <a href=\"https://www.kaggle.com/nomorevotch\" target=\"_blank\">sqrt4kaido</a>)</h1>\n<p>I've been participated in competitions as a solo, and this MLB competition was my first experience participating as a team. I learned a lot of things when I participated competitions by solo, but I learned even more this time. I could get a gold medal thank to your a lot of ideas.</p>\n<h1>tea's part</h1>\n<h2>Models</h2>\n<ul>\n<li>LGBM  lag / no lag</li>\n<li>CatBoost  lag / no lag<br>\nUse <strong>optuna</strong> to optimize hyper-parameters.</li>\n</ul>\n<h2>Train &amp; Valid</h2>\n<ul>\n<li>Use 2018-01-01 ~ 2021-05-31 data for train, and 2021-06-01 ~ 2021-07-17 for validation</li>\n<li>During the test period, I use 2018-01-01 ~ 2021-06-30 data for train</li>\n<li>Limit data to in-season data only</li>\n</ul>\n<h2>Features mainly effective</h2>\n<ul>\n<li>target lags (for 45 days)</li>\n<li>player target statistics (mean, median, max, min, var, skew, kurt)<ul>\n<li>Use the respective statistics for April, May, and June 2021.</li>\n<li>Also use the respective statistics for game day and no game day.</li></ul></li>\n<li>team target statistics (mean, median, max, min, var, skew, kurt)<ul>\n<li>Use statistics for June 2021 only</li></ul></li>\n<li>daysSinceLastGame / Roster</li>\n<li>days from the beginning of the year / month</li>\n<li>day of week</li>\n<li>years from the debut year</li>\n<li>age</li>\n<li>position</li>\n<li>player status</li>\n<li>playerBoxScores features<br>\n<br></li>\n</ul>\n<h1>sqrt4kaido part</h1>\n<h2>Models</h2>\n<ul>\n<li>LGBM only no lag</li>\n</ul>\n<ol>\n<li>The first seed, only the players in test are used for training.</li>\n<li>Second seed, only the players in test are used for training.</li>\n<li>The first seed. All the players are used for training.<br>\nUse <strong>optuna</strong> to optimize hyper-parameters.<br>\nBefore update train.csv, the best score for my single model was 1.3146.</li>\n</ol>\n<h2>Train &amp; Valid</h2>\n<ul>\n<li>training phase<br>\nUse 2018-01-01 ~ 2021-03-31 data for train, and 2021-04-01 ~ 2021-04-30 for validation</li>\n<li>after update train.csv<br>\nUse 2018-01-01 ~ 2021-05-31 data for train, and 2021-06-01 ~ 2021-06-30 for validation<br>\nI did not use the data for the first half of July because, unlike the other months, the values were very small.</li>\n<li>Limit data to in-season data only &lt;- important</li>\n</ul>\n<h2>Features mainly effective</h2>\n<ul>\n<li>player target statistics (mean, median, max, min, var, skew, kurt)<br>\nUse the respective statistics for June 2021. This feature is leaked to the valid data, but we used it because it was also valid for public testing.</li>\n<li>team target statistics (mean, median, max, min, var, skew, kurt)<br>\nsame as above</li>\n<li>season info<br>\npre/post season? regular season? alltardate? etc.</li>\n<li>award flag</li>\n<li>daysSinceLastGame / Roster &lt;- important</li>\n<li>day of week</li>\n<li>years from the debut year</li>\n<li>age</li>\n<li>position</li>\n<li>player status</li>\n<li>playerBoxScores features</li>\n</ul>\n<h2>did not work</h2>\n<ul>\n<li>Create separate models for the data with and without matches.<br>\nAfter update train.csv, it stopped working for some reason.</li>\n<li>Undo the scale and learn<br>\nThis time the target is scaled from 0 to 100, but there was a <a href=\"https://www.kaggle.com/c/mlb-player-digital-engagement-forecasting/discussion/253940\" target=\"_blank\">discussion</a> that it seems to be max-scaled for each day.<br>\nSo I tried training with the scale back, and the CV results improved. However, when we actually applied it to the test data, the score worsened because the wrong scaling was applied when the max prediction was shifted, so all the predictions for the day were shifted.<br>\nI also tried scaling to the average of the same month in the previous year, but this did not improve the accuracy.<br>\n<br></li>\n</ul>\n<h1>Makabe's part</h1>\n<h2>Models</h2>\n<ul>\n<li>I used three models.<ul>\n<li>Simple NN (add lags)</li>\n<li>Light GBM (add lags)</li>\n<li>Light GBM (none lags)</li></ul></li>\n<li>The reason to use three type models is the following.<ul>\n<li>To ensemble is useful making robust models.</li>\n<li>Regarding the lag feature, <strong>which is the key to this competition</strong>, I noticed the following fact in the middle of the competition.</li>\n<li>(If we can get the correct target information as a lag early in the test period.)<ul>\n<li>Light GBM (add lags) &gt; Simple NN (add lags) &gt; Light GBM (none lags)</li></ul></li>\n<li>(If we have to use the target information predicted by model as lag late in the test period.)<ul>\n<li>Light GBM (none lags) &gt; Simple NN (add lags) &gt; Light GBM (add lags)</li></ul></li>\n<li>From the above, I can see that lag feature is very effective if it contains the correct values. However, because of its effectiveness, there is also a possibility of overfitting. This tendency was especially for LGB.</li>\n<li>I decided to control the blend ratio of the three models according to the number of days in test period to balance lag features (effectiveness &amp; overfitting).</li>\n<li>The method of controlling the size of lag features depending on test time period was pointed out by <a href=\"https://www.kaggle.com/c/mlb-player-digital-engagement-forecasting/discussion/256620\" target=\"_blank\">other Kagglers</a> too. It was very effective method and we could this knowledge for our entire team ensemble too.</li></ul></li>\n</ul>\n<h2>Training / Validation</h2>\n<ul>\n<li>I used 5 folds.</li>\n</ul>\n<ol>\n<li>&lt; 20200801(training) &amp; 2020/08/01 - 2020/09/30(validation)</li>\n<li>&lt; 20200901(training) &amp; 2020/09/01 - 2020/10/31(validation)</li>\n<li>&lt; 20210401(training) &amp; 2021/04/01 - 2021/05/31(validation)</li>\n<li>&lt; 20210501(training) &amp; 2021/05/01 - 2021/06/30(validation)</li>\n<li>&lt; 20210601(training) &amp; 2021/06/01 - 2021/07/xx(validation)</li>\n</ol>\n<ul>\n<li>I decided on period with the following two points in mind.<ul>\n<li>Season with irregular schedule due to Covid-19 (2020).</li>\n<li>The trend of target feature is very different between the first and second half of the season.</li></ul></li>\n<li>However, I hadn't been able to determine the period above based on a deep consideration of this impact. I think there is room for reconsideration.</li>\n</ul>\n<h2>Features</h2>\n<p>The following is a list of feature what is helped improve accuracy. (There were many useful features than the below, but most of them were already in other Kaggler's notebooks, so I omit them.)</p>\n<ul>\n<li>stats of target feature<ul>\n<li>I labeled the data by period as follows, and I added stats that is the one before as features.</li>\n<li>2018.3 &amp; 2018.4 (=label1), 2018.5 &amp; 2018.6 (=label2), …</li>\n<li>I thought that by doing this, and I could use stats of target until September 2021. However, I knew after the competition, test period was until August 31, so I didn't need to use the stats of target two months earlier.</li>\n<li>I used mean, median, distribution,.. and more.</li>\n<li>I calculated the stats of targets by defense position, roster status, and team, respectively.</li>\n<li>There was difference of targets by roster status. (The difference was especially noticeable in target1.)</li></ul></li>\n<li>number of days what have elapsed since the last time player participated in game.<ul>\n<li>This ideas was from <a href=\"https://www.kaggle.com/nomorevotch\" target=\"_blank\">@sqrt4kaido</a>. This variable was based on the fact that the engagement of players who haven't played for a long time tends to decline over time, and contributed to improving accuracy.</li></ul></li>\n<li>score from the day before yesterday<ul>\n<li>I thought the success of a game is continuous, and player's popularity is not necessarily based only on the previous day's performance, but past success is a factor.</li>\n<li><a href=\"https://www.kaggle.com/c/mlb-player-digital-engagement-forecasting/discussion/266596\" target=\"_blank\">Team AutoMLB seems to this point</a>, and incorporated player scores into the model for a considerable period of time. I probably should have included more time periods than just the day before yesterday.</li></ul></li>\n<li>A lot of stats<ul>\n<li>I used various features based on <a href=\"https://sabr.org/sabermetrics\" target=\"_blank\">Sabermetric</a>. (K/BB, RC, OPS, FIP.. and more)</li>\n<li>This is especially for pitchers, I thought the ratio divided by the number of innings pitched or the number of pitches thrown tends to be a better representation of the player than the absolute number. (It's no surprise that starting pitchers strike out more than relief pitchers.) For this reason, I also incorporate many indicators divided by the number of innings or throw of pitches(K9, HR9.. and more).<br>\n<br></li></ul></li>\n</ul>\n<h1>Doing by team</h1>\n<ul>\n<li>In order to ensemble easily, it was necessary to improve the reproducibility and reusability of each member’s source code. Therefore, we unified the specifications and developed the prediction classes with common methods.</li>\n<li>In this competition, it was very important to finsh the process correctly, because the Time Series API is very complicated. Therefore, we created a LocalTestClass to reproduce the Time Series API in local environment.</li>\n<li>By using this class, we could verify the accuracy in a larger period.</li>\n<li>We controlled the ensemble ratio for the fact that the effectiveness of lag feature weakens over time.<br>\n<a href=\"https://postimg.cc/gwwHCDJm\" target=\"_blank\"><img src=\"https://i.postimg.cc/g0st6tbX/MLB-sub.png\" alt=\"MLB-sub.png\"></a></li>\n</ul>",
      "rawMarkdown": "First of all, I also would like to thank this competition host and all competitors. It has been about two years since I got my first silver medal in the APTOS competition, and I'm finally a Kaggle master. I'm very pleased now. I'm still new, but I'll continue to learn from Kaggle with a sincere attitude.\n I am also happy that I could get a gold medal in a competition related to my favorite sport what is baseball. Baseball may be a minor sport on the global stage, but it's very popular in Japan, and I watch NPB (Nippon Professional Baseball Organization) games every week. I wish this fascinating and strategic would become more popular around the world!\n\n# To member ([tea](https://www.kaggle.com/tea1013) & [sqrt4kaido](https://www.kaggle.com/nomorevotch))\n I've been participated in competitions as a solo, and this MLB competition was my first experience participating as a team. I learned a lot of things when I participated competitions by solo, but I learned even more this time. I could get a gold medal thank to your a lot of ideas.\n\n# tea's part\n## Models\n- LGBM  lag / no lag\n- CatBoost  lag / no lag\nUse **optuna** to optimize hyper-parameters.\n## Train & Valid\n- Use 2018-01-01 ~ 2021-05-31 data for train, and 2021-06-01 ~ 2021-07-17 for validation\n- During the test period, I use 2018-01-01 ~ 2021-06-30 data for train\n- Limit data to in-season data only\n## Features mainly effective\n- target lags (for 45 days)\n- player target statistics (mean, median, max, min, var, skew, kurt)\n    - Use the respective statistics for April, May, and June 2021.\n    - Also use the respective statistics for game day and no game day.\n- team target statistics (mean, median, max, min, var, skew, kurt)\n    - Use statistics for June 2021 only\n- daysSinceLastGame / Roster\n- days from the beginning of the year / month\n- day of week\n- years from the debut year\n- age\n- position\n- player status\n- playerBoxScores features\n\n<br />\n\n# sqrt4kaido part\n## Models\n- LGBM only no lag\n1. The first seed, only the players in test are used for training.\n2. Second seed, only the players in test are used for training.\n3. The first seed. All the players are used for training.\nUse **optuna** to optimize hyper-parameters.\nBefore update train.csv, the best score for my single model was 1.3146.\n## Train & Valid\n- training phase\nUse 2018-01-01 ~ 2021-03-31 data for train, and 2021-04-01 ~ 2021-04-30 for validation\n- after update train.csv\nUse 2018-01-01 ~ 2021-05-31 data for train, and 2021-06-01 ~ 2021-06-30 for validation\nI did not use the data for the first half of July because, unlike the other months, the values were very small.\n- Limit data to in-season data only <- important\n## Features mainly effective\n- player target statistics (mean, median, max, min, var, skew, kurt)\nUse the respective statistics for June 2021. This feature is leaked to the valid data, but we used it because it was also valid for public testing.\n- team target statistics (mean, median, max, min, var, skew, kurt)\nsame as above\n- season info\npre/post season? regular season? alltardate? etc.\n- award flag\n- daysSinceLastGame / Roster <- important\n- day of week\n- years from the debut year\n- age\n- position\n- player status\n- playerBoxScores features\n## did not work\n- Create separate models for the data with and without matches.\nAfter update train.csv, it stopped working for some reason.\n- Undo the scale and learn\nThis time the target is scaled from 0 to 100, but there was a [discussion](https://www.kaggle.com/c/mlb-player-digital-engagement-forecasting/discussion/253940) that it seems to be max-scaled for each day.\nSo I tried training with the scale back, and the CV results improved. However, when we actually applied it to the test data, the score worsened because the wrong scaling was applied when the max prediction was shifted, so all the predictions for the day were shifted.\nI also tried scaling to the average of the same month in the previous year, but this did not improve the accuracy.\n\n<br />\n\n# Makabe's part\n## Models\n- I used three models.\n  - Simple NN (add lags)\n  - Light GBM (add lags)\n  - Light GBM (none lags)\n\n- The reason to use three type models is the following.\n  - To ensemble is useful making robust models.\n  - Regarding the lag feature, <strong>which is the key to this competition</strong>, I noticed the following fact in the middle of the competition.\n    - (If we can get the correct target information as a lag early in the test period.)\n      -  Light GBM (add lags) > Simple NN (add lags) > Light GBM (none lags)\n    - (If we have to use the target information predicted by model as lag late in the test period.)\n      -  Light GBM (none lags) > Simple NN (add lags) > Light GBM (add lags)\n    - From the above, I can see that lag feature is very effective if it contains the correct values. However, because of its effectiveness, there is also a possibility of overfitting. This tendency was especially for LGB.\n    - I decided to control the blend ratio of the three models according to the number of days in test period to balance lag features (effectiveness & overfitting).\n    - The method of controlling the size of lag features depending on test time period was pointed out by [other Kagglers](https://www.kaggle.com/c/mlb-player-digital-engagement-forecasting/discussion/256620) too. <string>It was very effective method</strong> and we could this knowledge for our entire team ensemble too.\n\n## Training / Validation\n- I used 5 folds.\n  1. < 20200801(training) & 2020/08/01 - 2020/09/30(validation)\n  2. < 20200901(training) & 2020/09/01 - 2020/10/31(validation)\n  3. < 20210401(training) & 2021/04/01 - 2021/05/31(validation)\n  4. < 20210501(training) & 2021/05/01 - 2021/06/30(validation)\n  5. < 20210601(training) & 2021/06/01 - 2021/07/xx(validation)\n\n- I decided on period with the following two points in mind.\n  - Season with irregular schedule due to Covid-19 (2020).\n  - The trend of target feature is very different between the first and second half of the season.\n\n- However, I hadn't been able to determine the period above based on a deep consideration of this impact. I think there is room for reconsideration.\n\n## Features\n The following is a list of feature what is helped improve accuracy. (There were many useful features than the below, but most of them were already in other Kaggler's notebooks, so I omit them.)\n- stats of target feature\n  - I labeled the data by period as follows, and I added stats that is the one before as features.\n    - 2018.3 & 2018.4 (=label1), 2018.5 & 2018.6 (=label2), ...\n    - I thought that by doing this, and I could use stats of target until September 2021. However, I knew after the competition, test period was until August 31, so I didn't need to use the stats of target two months earlier.\n  - I used mean, median, distribution,.. and more.\n  - I calculated the stats of targets by defense position, roster status, and team, respectively.\n    - There was difference of targets by roster status. (The difference was especially noticeable in target1.)\n\n- number of days what have elapsed since the last time player participated in game.\n  - This ideas was from [@sqrt4kaido](https://www.kaggle.com/nomorevotch). This variable was based on the fact that the engagement of players who haven't played for a long time tends to decline over time, and contributed to improving accuracy.\n\n- score from the day before yesterday\n  - I thought the success of a game is continuous, and player's popularity is not necessarily based only on the previous day's performance, but past success is a factor.\n  - [Team AutoMLB seems to this point](https://www.kaggle.com/c/mlb-player-digital-engagement-forecasting/discussion/266596), and incorporated player scores into the model for a considerable period of time. I probably should have included more time periods than just the day before yesterday.\n\n- A lot of stats\n  - I used various features based on [Sabermetric](https://sabr.org/sabermetrics). (K/BB, RC, OPS, FIP.. and more)\n  - This is especially for pitchers, I thought the ratio divided by the number of innings pitched or the number of pitches thrown tends to be a better representation of the player than the absolute number. (It's no surprise that starting pitchers strike out more than relief pitchers.) For this reason, I also incorporate many indicators divided by the number of innings or throw of pitches(K9, HR9.. and more).\n\n<br />\n\n# Doing by team\n- In order to ensemble easily, it was necessary to improve the reproducibility and reusability of each member’s source code. Therefore, we unified the specifications and developed the prediction classes with common methods.\n- In this competition, it was very important to finsh the process correctly, because the Time Series API is very complicated. Therefore, we created a LocalTestClass to reproduce the Time Series API in local environment.\n- By using this class, we could verify the accuracy in a larger period.\n- We controlled the ensemble ratio for the fact that the effectiveness of lag feature weakens over time.\n\n[![MLB-sub.png](https://i.postimg.cc/g0st6tbX/MLB-sub.png)](https://postimg.cc/gwwHCDJm)",
      "votes": null
    },
    {
      "id": "1511403",
      "postDate": "09/13/2021 11:49:08",
      "content": "<p>Congratulations! In this competition, it is a very difficult to build such a complex data pipeline with a team. Do you have any tips on how to make it work?</p>",
      "rawMarkdown": "Congratulations! In this competition, it is a very difficult to build such a complex data pipeline with a team. Do you have any tips on how to make it work?",
      "votes": null
    },
    {
      "id": "1511520",
      "postDate": "09/13/2021 13:46:04",
      "content": "<p><a href=\"https://www.kaggle.com/nyanpn\" target=\"_blank\">@nyanpn</a> </p>\n<p>Thank you, congratulations to you too! Your solution was very interesting for me.<br>\n Regarding this competition, Implementing a source code is very complicated as you pointed out..<br>\n We decided to create classes with a common rules, and restricted using global area.</p>\n<p>I list some typical rules below.</p>\n<ul>\n<li>The Global area must be limited to the part of bellow.<ol>\n<li>loading data</li></ol><ul>\n<li>We could pprevent mistakes to use common table.</li>\n<li>We could get the correct target by 31 July.</li></ul><ol>\n<li>instantiating models</li>\n<li>preprocess</li>\n<li>inference</li></ol></li>\n<li>We created methods with simple rules. The following is image.</li>\n</ul>\n<pre><code>data = read_data()\n\nc1 = Class1()\nc2 = Class2()\nc3 = Class3()\n\n# cumulative features (targetLag, yesterdayScore, sinceLastGames, ..)\ndata1 = c1.preprocess(data)\ndata2 = c2.preprocess(data)\ndata3 = c3.preprocess(data)\n\nfor test_df, sample in iter_test:\n    p1, data1 = c1.pred(test_df, sample, data1)\n    p2, data2 = c2.pred(test_df, sample, data2)\n    p3, data3 = c3.pred(test_df, sample, data3)\n\n    p_final = p1 * x1 p2 * x2 + p3 * x3\n</code></pre>\n<ul>\n<li>Probably \"cumulative features\" are the most of difficult part in this competition, and I think many people made mistakes when they updated this datas.</li>\n<li>tea and kaido choose to keep it as classes variables, and I decided to update at global area. (e.g. above)</li>\n<li>In any case, we made many mistakes..</li>\n<li>I think it was good that we were checked the behavior in various patterns (earlier or later start date of the test period, missing data, ..etc.) finally.</li>\n</ul>",
      "rawMarkdown": "nyanpn \n\n Thank you, congratulations to you too! Your solution was very interesting for me.\n Regarding this competition, Implementing a source code is very complicated as you pointed out..\n We decided to create classes with a common rules, and restricted using global area.\n\n I list some typical rules below.\n\n- The Global area must be limited to the part of bellow.\n  1. loading data\n    - We could pprevent mistakes to use common table.\n    - We could get the correct target by 31 July.\n  2. instantiating models\n  3. preprocess\n  4. inference\n- We created methods with simple rules. The following is image.\n\n\n```\ndata = read_data()\n\nc1 = Class1()\nc2 = Class2()\nc3 = Class3()\n\n# cumulative features (targetLag, yesterdayScore, sinceLastGames, ..)\ndata1 = c1.preprocess(data)\ndata2 = c2.preprocess(data)\ndata3 = c3.preprocess(data)\n\nfor test_df, sample in iter_test:\n    p1, data1 = c1.pred(test_df, sample, data1)\n    p2, data2 = c2.pred(test_df, sample, data2)\n    p3, data3 = c3.pred(test_df, sample, data3)\n\n    p_final = p1 * x1 p2 * x2 + p3 * x3\n```\n\n- Probably \"cumulative features\" are the most of difficult part in this competition, and I think many people made mistakes when they updated this datas.\n- tea and kaido choose to keep it as classes variables, and I decided to update at global area. (e.g. above)\n- In any case, we made many mistakes..\n- I think it was good that we were checked the behavior in various patterns (earlier or later start date of the test period, missing data, ..etc.) finally.",
      "votes": null
    },
    {
      "id": "1512073",
      "postDate": "09/13/2021 23:16:25",
      "content": "<p>Thanks! I think it is a very good discipline to encapsulate the code of each member in a class.</p>",
      "rawMarkdown": "Thanks! I think it is a very good discipline to encapsulate the code of each member in a class.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1511403,
      "author_name": "nyanpn",
      "author_url": "",
      "post_date": "09/13/2021 11:49:08",
      "content": "<p>Congratulations! In this competition, it is a very difficult to build such a complex data pipeline with a team. Do you have any tips on how to make it work?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1511520,
      "author_name": "spidermandance",
      "author_url": "",
      "post_date": "09/13/2021 13:46:04",
      "content": "<p><a href=\"https://www.kaggle.com/nyanpn\" target=\"_blank\">@nyanpn</a> </p>\n<p>Thank you, congratulations to you too! Your solution was very interesting for me.<br>\n Regarding this competition, Implementing a source code is very complicated as you pointed out..<br>\n We decided to create classes with a common rules, and restricted using global area.</p>\n<p>I list some typical rules below.</p>\n<ul>\n<li>The Global area must be limited to the part of bellow.<ol>\n<li>loading data</li></ol><ul>\n<li>We could pprevent mistakes to use common table.</li>\n<li>We could get the correct target by 31 July.</li></ul><ol>\n<li>instantiating models</li>\n<li>preprocess</li>\n<li>inference</li></ol></li>\n<li>We created methods with simple rules. The following is image.</li>\n</ul>\n<pre><code>data = read_data()\n\nc1 = Class1()\nc2 = Class2()\nc3 = Class3()\n\n# cumulative features (targetLag, yesterdayScore, sinceLastGames, ..)\ndata1 = c1.preprocess(data)\ndata2 = c2.preprocess(data)\ndata3 = c3.preprocess(data)\n\nfor test_df, sample in iter_test:\n    p1, data1 = c1.pred(test_df, sample, data1)\n    p2, data2 = c2.pred(test_df, sample, data2)\n    p3, data3 = c3.pred(test_df, sample, data3)\n\n    p_final = p1 * x1 p2 * x2 + p3 * x3\n</code></pre>\n<ul>\n<li>Probably \"cumulative features\" are the most of difficult part in this competition, and I think many people made mistakes when they updated this datas.</li>\n<li>tea and kaido choose to keep it as classes variables, and I decided to update at global area. (e.g. above)</li>\n<li>In any case, we made many mistakes..</li>\n<li>I think it was good that we were checked the behavior in various patterns (earlier or later start date of the test period, missing data, ..etc.) finally.</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1512073,
          "author_name": "nyanpn",
          "author_url": "",
          "post_date": "09/13/2021 23:16:25",
          "content": "<p>Thanks! I think it is a very good discipline to encapsulate the code of each member in a class.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1511022": "First of all, I also would like to thank this competition host and all competitors. It has been about two years since I got my first silver medal in the APTOS competition, and I'm finally a Kaggle master. I'm very pleased now. I'm still new, but I'll continue to learn from Kaggle with a sincere attitude.\n I am also happy that I could get a gold medal in a competition related to my favorite sport what is baseball. Baseball may be a minor sport on the global stage, but it's very popular in Japan, and I watch NPB (Nippon Professional Baseball Organization) games every week. I wish this fascinating and strategic would become more popular around the world!\n\n# To member ([tea](https://www.kaggle.com/tea1013) & [sqrt4kaido](https://www.kaggle.com/nomorevotch))\n I've been participated in competitions as a solo, and this MLB competition was my first experience participating as a team. I learned a lot of things when I participated competitions by solo, but I learned even more this time. I could get a gold medal thank to your a lot of ideas.\n\n# tea's part\n## Models\n- LGBM  lag / no lag\n- CatBoost  lag / no lag\nUse **optuna** to optimize hyper-parameters.\n## Train & Valid\n- Use 2018-01-01 ~ 2021-05-31 data for train, and 2021-06-01 ~ 2021-07-17 for validation\n- During the test period, I use 2018-01-01 ~ 2021-06-30 data for train\n- Limit data to in-season data only\n## Features mainly effective\n- target lags (for 45 days)\n- player target statistics (mean, median, max, min, var, skew, kurt)\n    - Use the respective statistics for April, May, and June 2021.\n    - Also use the respective statistics for game day and no game day.\n- team target statistics (mean, median, max, min, var, skew, kurt)\n    - Use statistics for June 2021 only\n- daysSinceLastGame / Roster\n- days from the beginning of the year / month\n- day of week\n- years from the debut year\n- age\n- position\n- player status\n- playerBoxScores features\n\n<br />\n\n# sqrt4kaido part\n## Models\n- LGBM only no lag\n1. The first seed, only the players in test are used for training.\n2. Second seed, only the players in test are used for training.\n3. The first seed. All the players are used for training.\nUse **optuna** to optimize hyper-parameters.\nBefore update train.csv, the best score for my single model was 1.3146.\n## Train & Valid\n- training phase\nUse 2018-01-01 ~ 2021-03-31 data for train, and 2021-04-01 ~ 2021-04-30 for validation\n- after update train.csv\nUse 2018-01-01 ~ 2021-05-31 data for train, and 2021-06-01 ~ 2021-06-30 for validation\nI did not use the data for the first half of July because, unlike the other months, the values were very small.\n- Limit data to in-season data only <- important\n## Features mainly effective\n- player target statistics (mean, median, max, min, var, skew, kurt)\nUse the respective statistics for June 2021. This feature is leaked to the valid data, but we used it because it was also valid for public testing.\n- team target statistics (mean, median, max, min, var, skew, kurt)\nsame as above\n- season info\npre/post season? regular season? alltardate? etc.\n- award flag\n- daysSinceLastGame / Roster <- important\n- day of week\n- years from the debut year\n- age\n- position\n- player status\n- playerBoxScores features\n## did not work\n- Create separate models for the data with and without matches.\nAfter update train.csv, it stopped working for some reason.\n- Undo the scale and learn\nThis time the target is scaled from 0 to 100, but there was a [discussion](https://www.kaggle.com/c/mlb-player-digital-engagement-forecasting/discussion/253940) that it seems to be max-scaled for each day.\nSo I tried training with the scale back, and the CV results improved. However, when we actually applied it to the test data, the score worsened because the wrong scaling was applied when the max prediction was shifted, so all the predictions for the day were shifted.\nI also tried scaling to the average of the same month in the previous year, but this did not improve the accuracy.\n\n<br />\n\n# Makabe's part\n## Models\n- I used three models.\n  - Simple NN (add lags)\n  - Light GBM (add lags)\n  - Light GBM (none lags)\n\n- The reason to use three type models is the following.\n  - To ensemble is useful making robust models.\n  - Regarding the lag feature, <strong>which is the key to this competition</strong>, I noticed the following fact in the middle of the competition.\n    - (If we can get the correct target information as a lag early in the test period.)\n      -  Light GBM (add lags) > Simple NN (add lags) > Light GBM (none lags)\n    - (If we have to use the target information predicted by model as lag late in the test period.)\n      -  Light GBM (none lags) > Simple NN (add lags) > Light GBM (add lags)\n    - From the above, I can see that lag feature is very effective if it contains the correct values. However, because of its effectiveness, there is also a possibility of overfitting. This tendency was especially for LGB.\n    - I decided to control the blend ratio of the three models according to the number of days in test period to balance lag features (effectiveness & overfitting).\n    - The method of controlling the size of lag features depending on test time period was pointed out by [other Kagglers](https://www.kaggle.com/c/mlb-player-digital-engagement-forecasting/discussion/256620) too. <string>It was very effective method</strong> and we could this knowledge for our entire team ensemble too.\n\n## Training / Validation\n- I used 5 folds.\n  1. < 20200801(training) & 2020/08/01 - 2020/09/30(validation)\n  2. < 20200901(training) & 2020/09/01 - 2020/10/31(validation)\n  3. < 20210401(training) & 2021/04/01 - 2021/05/31(validation)\n  4. < 20210501(training) & 2021/05/01 - 2021/06/30(validation)\n  5. < 20210601(training) & 2021/06/01 - 2021/07/xx(validation)\n\n- I decided on period with the following two points in mind.\n  - Season with irregular schedule due to Covid-19 (2020).\n  - The trend of target feature is very different between the first and second half of the season.\n\n- However, I hadn't been able to determine the period above based on a deep consideration of this impact. I think there is room for reconsideration.\n\n## Features\n The following is a list of feature what is helped improve accuracy. (There were many useful features than the below, but most of them were already in other Kaggler's notebooks, so I omit them.)\n- stats of target feature\n  - I labeled the data by period as follows, and I added stats that is the one before as features.\n    - 2018.3 & 2018.4 (=label1), 2018.5 & 2018.6 (=label2), ...\n    - I thought that by doing this, and I could use stats of target until September 2021. However, I knew after the competition, test period was until August 31, so I didn't need to use the stats of target two months earlier.\n  - I used mean, median, distribution,.. and more.\n  - I calculated the stats of targets by defense position, roster status, and team, respectively.\n    - There was difference of targets by roster status. (The difference was especially noticeable in target1.)\n\n- number of days what have elapsed since the last time player participated in game.\n  - This ideas was from [@sqrt4kaido](https://www.kaggle.com/nomorevotch). This variable was based on the fact that the engagement of players who haven't played for a long time tends to decline over time, and contributed to improving accuracy.\n\n- score from the day before yesterday\n  - I thought the success of a game is continuous, and player's popularity is not necessarily based only on the previous day's performance, but past success is a factor.\n  - [Team AutoMLB seems to this point](https://www.kaggle.com/c/mlb-player-digital-engagement-forecasting/discussion/266596), and incorporated player scores into the model for a considerable period of time. I probably should have included more time periods than just the day before yesterday.\n\n- A lot of stats\n  - I used various features based on [Sabermetric](https://sabr.org/sabermetrics). (K/BB, RC, OPS, FIP.. and more)\n  - This is especially for pitchers, I thought the ratio divided by the number of innings pitched or the number of pitches thrown tends to be a better representation of the player than the absolute number. (It's no surprise that starting pitchers strike out more than relief pitchers.) For this reason, I also incorporate many indicators divided by the number of innings or throw of pitches(K9, HR9.. and more).\n\n<br />\n\n# Doing by team\n- In order to ensemble easily, it was necessary to improve the reproducibility and reusability of each member’s source code. Therefore, we unified the specifications and developed the prediction classes with common methods.\n- In this competition, it was very important to finsh the process correctly, because the Time Series API is very complicated. Therefore, we created a LocalTestClass to reproduce the Time Series API in local environment.\n- By using this class, we could verify the accuracy in a larger period.\n- We controlled the ensemble ratio for the fact that the effectiveness of lag feature weakens over time.\n\n[![MLB-sub.png](https://i.postimg.cc/g0st6tbX/MLB-sub.png)](https://postimg.cc/gwwHCDJm)",
    "1511403": "Congratulations! In this competition, it is a very difficult to build such a complex data pipeline with a team. Do you have any tips on how to make it work?",
    "1511520": "nyanpn \n\n Thank you, congratulations to you too! Your solution was very interesting for me.\n Regarding this competition, Implementing a source code is very complicated as you pointed out..\n We decided to create classes with a common rules, and restricted using global area.\n\n I list some typical rules below.\n\n- The Global area must be limited to the part of bellow.\n  1. loading data\n    - We could pprevent mistakes to use common table.\n    - We could get the correct target by 31 July.\n  2. instantiating models\n  3. preprocess\n  4. inference\n- We created methods with simple rules. The following is image.\n\n\n```\ndata = read_data()\n\nc1 = Class1()\nc2 = Class2()\nc3 = Class3()\n\n# cumulative features (targetLag, yesterdayScore, sinceLastGames, ..)\ndata1 = c1.preprocess(data)\ndata2 = c2.preprocess(data)\ndata3 = c3.preprocess(data)\n\nfor test_df, sample in iter_test:\n    p1, data1 = c1.pred(test_df, sample, data1)\n    p2, data2 = c2.pred(test_df, sample, data2)\n    p3, data3 = c3.pred(test_df, sample, data3)\n\n    p_final = p1 * x1 p2 * x2 + p3 * x3\n```\n\n- Probably \"cumulative features\" are the most of difficult part in this competition, and I think many people made mistakes when they updated this datas.\n- tea and kaido choose to keep it as classes variables, and I decided to update at global area. (e.g. above)\n- In any case, we made many mistakes..\n- I think it was good that we were checked the behavior in various patterns (earlier or later start date of the test period, missing data, ..etc.) finally.",
    "1512073": "Thanks! I think it is a very good discipline to encapsulate the code of each member in a class."
  },
  "source": "meta"
}