{
  "id": 264396,
  "title": "11th place solution",
  "url": "/competitions/mlb-player-digital-engagement-forecasting/discussion/264396",
  "author_name": "Makotu",
  "post_date": "2021-08-12T03:05:51.223000",
  "votes": 25,
  "comment_count": 0,
  "views": 0,
  "content": "<p><strong>[update]</strong><br>\nIn the end, I got 11th place and I was able to stay within the gold medal range. <br>\nI'd like to thank the people running this competition and all the people I competed with! Thank you very much.</p>\n<p>a first week of RE-RUN is over. <br>\nI was lucky enough to have a successful rerun.<br>\nI would like to thank the admin and all the participants for organizing this event.<br>\nThere is still a month to go before it ends, but before I forget, I'll describe my solution.<br>\n(I'm not very good at English, so I basically translate Japanese at DeepL. I apologize if you have trouble reading.)<br>\n<br></p>\n<hr>\n<h4>Difficult points of this competition</h4>\n<ol>\n<li>The target is scaled from 0 to 100 on a daily basis.</li>\n<li>There is a difference in the information available to pitchers and batters.(The roles of the pitcher and batter are fundamentally different.)</li>\n</ol>\n<p><br></p>\n<h4>My approach to the above points</h4>\n<ul>\n<li><p><strong>Use only in-season data for training</strong><br>\nWhen I checked the daily median values, I found that there was a considerable difference in the median values between the on-season and off-season.<br>\nI thought that on-season engagement would be determined by performance and success during the season, while off-season engagement would be determined by media exposure and social networking regardless of the season(and we can't use this kind of data).<br>\nIn this competition, we predicts the engagement during the season, so I judged that data from the off-season would be noisy and did not use it for training and validation.</p></li>\n<li><p><strong>Create features that represent what happened on any given day and the game</strong><br>\nI thought that daily median engagement scores might differ depending on what happened in the game as a whole and what happened in other games on that day.<br>\nTherefore, I created the features related to the actions that occurred throughout the day.</p>\n<ul>\n<li><p>Number of games played that day, total runs scored, home runs, hits, etc.</p></li>\n<li><p>The average pct of both teams in each game (whether there was an exciting game on the same day between teams with high winning percentages)</p></li>\n<li><p>The total twitter followers of the teams playing in the game (home and away) as a percentage of the total followers of all teams (whether the game was played between teams that are popular on social networking sites)</p></li>\n<li><p>The total number of twitter followers of the players who played in the game that day as a percentage of the total number of followers of all players (how influential are the players who played that day on SNS)</p></li>\n<li><p>How long has the season elapsed (what percentage of the schedule has been completed)?<br><br>\nIn addition to the above, I added the player's performance.<br><br></p></li>\n<li><p>Daily performance</p></li>\n<li><p>Season performance</p></li>\n<li><p>Previous season performance (last season's final performance)</p></li></ul></li>\n<li><p><strong>Built three different models</strong><br>\none for pitchers, one for batters, and one according to no play for the 2021 season.<br>\nThe features used in each model is slightly different from each other. <br>\nThe features for any given day or game described above are the same for all models, <br>\nbut in the pitcher's model, the only feature used in the player's performance was the performance related to pitching. in the batter's model, only the performance related to batting was used. <br>\nFor the model of a player without a season of play, we built the model without using any features related to performance.<br><br>\nThe reason for this, as mentioned earlier, is that pitchers and batters have distinctly different roles. For example, even if a pitcher gets three hits, his reputation will not be high if he is scored on by his crucial pitches.<br>\nIn addition, for players who did not play in the season, we could not obtain their performance in the current season at all and had no material to judge them, so we built a separate model with fewer features.(no use performance features) <br>\n<br><br>\nAs a side note, I don't use target-lag-feature.<br>\nWhen I was testing the three models separately, the accuracy of the pitcher and batter models decreased when using past targets, so I decided not to use variables related to past targets this competition.</p></li>\n</ul>\n<p><br></p>\n<h4>validation scheme</h4>\n<p>For training and validation,  I used data from each season.<br>\nIn addition, I did not use the data for 2020 because the number of games was small and we thought that there might be abnormal movements due to the corona disaster.<br>\nAll models are made with lightgbm.</p>\n<p><br></p>\n<h5>Before the train data was updated (LB:1.3104(Previous leaderboard))</h5>\n<ul>\n<li>Training : 4/1/2019 to 4/30/2019</li>\n<li>Validation : 5/1/2019 to 5/30/2019 → check accuracy and obtained parameters in this period</li>\n<li>Re-training: re-train with parameters obtained in the above period on data from 4/1/2021 to 4/30/2021</li>\n</ul>\n<p><br></p>\n<h5>After the train data is updated</h5>\n<p>After the train data was updated, I tested both.<br>\na. whether August of the previous regular season could be predicted well (2019. except for 2020) <br>\nb. whether the last month of the same season could be predicted well.<br>\n The parameters obtained for each were used to make each prediction, and the final submission was weighted  a and b.</p>\n<ul>\n<li><p>Validation a: </p>\n<ul>\n<li>Training : 4/1/2019 to 7/30/2019 </li>\n<li>Validation : 8/1/2019 to 8/31/2019</li>\n<li>check accuracy and obtaining parameters in this period</li></ul></li>\n<li><p>Validation b: </p>\n<ul>\n<li>Training : 4/1/2021 to 6/17/2021</li>\n<li>Validation : 6/18/2021 to 7/17/2021</li>\n<li>check accuracy and obtaining parameters in this period</li></ul></li>\n<li><p>Re-training: <br>\nRe-train on data from 2021/4/1 to 7/17 with parameters obtained in the above period, respectively.</p></li>\n</ul>\n<p><br></p>\n<p>Final sub 1: <br>\nPredicted value of the model made with the parameters obtained in verification a × 0.7 +<br>\nPredicted value of the model made with the parameters obtained in verification b × 0.3</p>\n<p>Final sub 2:<br>\nPredicted value of the model created with the parameters obtained in verification a × 0.3 + Predicted value of the model created with the parameters obtained in verification b × 0.7</p>\n<p><br></p>\n<hr>\n<h5>postscript</h5>\n<p>Since my solution does not use lag-target, I believe that it is inferior to models that use lag-target for training and inference if we only look at the results of the first week.<br>\nHowever, I imagine that the accuracy of these models using lag-targets may gradually decrease in future re-runs, and I expect that there will be a few shake-up/down.<br>\n(But first, I need to get the next rerun running correctly…)</p>\n<p>Thank you for reading this far!</p>",
  "messages": [
    {
      "id": 1467507,
      "postDate": "2021-08-12T03:05:51.223Z",
      "content": "<p><strong>[update]</strong><br>\nIn the end, I got 11th place and I was able to stay within the gold medal range. <br>\nI'd like to thank the people running this competition and all the people I competed with! Thank you very much.</p>\n<p>a first week of RE-RUN is over. <br>\nI was lucky enough to have a successful rerun.<br>\nI would like to thank the admin and all the participants for organizing this event.<br>\nThere is still a month to go before it ends, but before I forget, I'll describe my solution.<br>\n(I'm not very good at English, so I basically translate Japanese at DeepL. I apologize if you have trouble reading.)<br>\n<br></p>\n<hr>\n<h4>Difficult points of this competition</h4>\n<ol>\n<li>The target is scaled from 0 to 100 on a daily basis.</li>\n<li>There is a difference in the information available to pitchers and batters.(The roles of the pitcher and batter are fundamentally different.)</li>\n</ol>\n<p><br></p>\n<h4>My approach to the above points</h4>\n<ul>\n<li><p><strong>Use only in-season data for training</strong><br>\nWhen I checked the daily median values, I found that there was a considerable difference in the median values between the on-season and off-season.<br>\nI thought that on-season engagement would be determined by performance and success during the season, while off-season engagement would be determined by media exposure and social networking regardless of the season(and we can't use this kind of data).<br>\nIn this competition, we predicts the engagement during the season, so I judged that data from the off-season would be noisy and did not use it for training and validation.</p></li>\n<li><p><strong>Create features that represent what happened on any given day and the game</strong><br>\nI thought that daily median engagement scores might differ depending on what happened in the game as a whole and what happened in other games on that day.<br>\nTherefore, I created the features related to the actions that occurred throughout the day.</p>\n<ul>\n<li><p>Number of games played that day, total runs scored, home runs, hits, etc.</p></li>\n<li><p>The average pct of both teams in each game (whether there was an exciting game on the same day between teams with high winning percentages)</p></li>\n<li><p>The total twitter followers of the teams playing in the game (home and away) as a percentage of the total followers of all teams (whether the game was played between teams that are popular on social networking sites)</p></li>\n<li><p>The total number of twitter followers of the players who played in the game that day as a percentage of the total number of followers of all players (how influential are the players who played that day on SNS)</p></li>\n<li><p>How long has the season elapsed (what percentage of the schedule has been completed)?<br><br>\nIn addition to the above, I added the player's performance.<br><br></p></li>\n<li><p>Daily performance</p></li>\n<li><p>Season performance</p></li>\n<li><p>Previous season performance (last season's final performance)</p></li></ul></li>\n<li><p><strong>Built three different models</strong><br>\none for pitchers, one for batters, and one according to no play for the 2021 season.<br>\nThe features used in each model is slightly different from each other. <br>\nThe features for any given day or game described above are the same for all models, <br>\nbut in the pitcher's model, the only feature used in the player's performance was the performance related to pitching. in the batter's model, only the performance related to batting was used. <br>\nFor the model of a player without a season of play, we built the model without using any features related to performance.<br><br>\nThe reason for this, as mentioned earlier, is that pitchers and batters have distinctly different roles. For example, even if a pitcher gets three hits, his reputation will not be high if he is scored on by his crucial pitches.<br>\nIn addition, for players who did not play in the season, we could not obtain their performance in the current season at all and had no material to judge them, so we built a separate model with fewer features.(no use performance features) <br>\n<br><br>\nAs a side note, I don't use target-lag-feature.<br>\nWhen I was testing the three models separately, the accuracy of the pitcher and batter models decreased when using past targets, so I decided not to use variables related to past targets this competition.</p></li>\n</ul>\n<p><br></p>\n<h4>validation scheme</h4>\n<p>For training and validation,  I used data from each season.<br>\nIn addition, I did not use the data for 2020 because the number of games was small and we thought that there might be abnormal movements due to the corona disaster.<br>\nAll models are made with lightgbm.</p>\n<p><br></p>\n<h5>Before the train data was updated (LB:1.3104(Previous leaderboard))</h5>\n<ul>\n<li>Training : 4/1/2019 to 4/30/2019</li>\n<li>Validation : 5/1/2019 to 5/30/2019 → check accuracy and obtained parameters in this period</li>\n<li>Re-training: re-train with parameters obtained in the above period on data from 4/1/2021 to 4/30/2021</li>\n</ul>\n<p><br></p>\n<h5>After the train data is updated</h5>\n<p>After the train data was updated, I tested both.<br>\na. whether August of the previous regular season could be predicted well (2019. except for 2020) <br>\nb. whether the last month of the same season could be predicted well.<br>\n The parameters obtained for each were used to make each prediction, and the final submission was weighted  a and b.</p>\n<ul>\n<li><p>Validation a: </p>\n<ul>\n<li>Training : 4/1/2019 to 7/30/2019 </li>\n<li>Validation : 8/1/2019 to 8/31/2019</li>\n<li>check accuracy and obtaining parameters in this period</li></ul></li>\n<li><p>Validation b: </p>\n<ul>\n<li>Training : 4/1/2021 to 6/17/2021</li>\n<li>Validation : 6/18/2021 to 7/17/2021</li>\n<li>check accuracy and obtaining parameters in this period</li></ul></li>\n<li><p>Re-training: <br>\nRe-train on data from 2021/4/1 to 7/17 with parameters obtained in the above period, respectively.</p></li>\n</ul>\n<p><br></p>\n<p>Final sub 1: <br>\nPredicted value of the model made with the parameters obtained in verification a × 0.7 +<br>\nPredicted value of the model made with the parameters obtained in verification b × 0.3</p>\n<p>Final sub 2:<br>\nPredicted value of the model created with the parameters obtained in verification a × 0.3 + Predicted value of the model created with the parameters obtained in verification b × 0.7</p>\n<p><br></p>\n<hr>\n<h5>postscript</h5>\n<p>Since my solution does not use lag-target, I believe that it is inferior to models that use lag-target for training and inference if we only look at the results of the first week.<br>\nHowever, I imagine that the accuracy of these models using lag-targets may gradually decrease in future re-runs, and I expect that there will be a few shake-up/down.<br>\n(But first, I need to get the next rerun running correctly…)</p>\n<p>Thank you for reading this far!</p>",
      "rawMarkdown": "**[update]**\nIn the end, I got 11th place and I was able to stay within the gold medal range. \nI'd like to thank the people running this competition and all the people I competed with! Thank you very much.\n\n\na first week of RE-RUN is over. \nI was lucky enough to have a successful rerun.\nI would like to thank the admin and all the participants for organizing this event.\nThere is still a month to go before it ends, but before I forget, I'll describe my solution.\n(I'm not very good at English, so I basically translate Japanese at DeepL. I apologize if you have trouble reading.)\n<br>\n\n\n***\n\n\n####  Difficult points of this competition\n1. The target is scaled from 0 to 100 on a daily basis.\n1. There is a difference in the information available to pitchers and batters.(The roles of the pitcher and batter are fundamentally different.)\n\n<br>\n\n#### My approach to the above points\n- **Use only in-season data for training**\nWhen I checked the daily median values, I found that there was a considerable difference in the median values between the on-season and off-season.\nI thought that on-season engagement would be determined by performance and success during the season, while off-season engagement would be determined by media exposure and social networking regardless of the season(and we can't use this kind of data).\nIn this competition, we predicts the engagement during the season, so I judged that data from the off-season would be noisy and did not use it for training and validation.\n\n- **Create features that represent what happened on any given day and the game**\nI thought that daily median engagement scores might differ depending on what happened in the game as a whole and what happened in other games on that day.\nTherefore, I created the features related to the actions that occurred throughout the day.\n  - Number of games played that day, total runs scored, home runs, hits, etc.\n  - The average pct of both teams in each game (whether there was an exciting game on the same day between teams with high winning percentages)\n  - The total twitter followers of the teams playing in the game (home and away) as a percentage of the total followers of all teams (whether the game was played between teams that are popular on social networking sites)\n  - The total number of twitter followers of the players who played in the game that day as a percentage of the total number of followers of all players (how influential are the players who played that day on SNS)\n  - How long has the season elapsed (what percentage of the schedule has been completed)?<br>\nIn addition to the above, I added the player's performance.<br><br>\n\n  - Daily performance\n  - Season performance\n  - Previous season performance (last season's final performance)\n\n- **Built three different models**\none for pitchers, one for batters, and one according to no play for the 2021 season.\nThe features used in each model is slightly different from each other. \nThe features for any given day or game described above are the same for all models, \nbut in the pitcher's model, the only feature used in the player's performance was the performance related to pitching. in the batter's model, only the performance related to batting was used. \nFor the model of a player without a season of play, we built the model without using any features related to performance.<br>\nThe reason for this, as mentioned earlier, is that pitchers and batters have distinctly different roles. For example, even if a pitcher gets three hits, his reputation will not be high if he is scored on by his crucial pitches.\nIn addition, for players who did not play in the season, we could not obtain their performance in the current season at all and had no material to judge them, so we built a separate model with fewer features.(no use performance features) \n<br>\nAs a side note, I don't use target-lag-feature.\nWhen I was testing the three models separately, the accuracy of the pitcher and batter models decreased when using past targets, so I decided not to use variables related to past targets this competition.\n\n\n<br>\n\n#### validation scheme\nFor training and validation,  I used data from each season.\nIn addition, I did not use the data for 2020 because the number of games was small and we thought that there might be abnormal movements due to the corona disaster.\nAll models are made with lightgbm.\n\n<br>\n\n##### Before the train data was updated (LB:1.3104(Previous leaderboard))\n - Training : 4/1/2019 to 4/30/2019\n - Validation : 5/1/2019 to 5/30/2019 → check accuracy and obtained parameters in this period\n - Re-training: re-train with parameters obtained in the above period on data from 4/1/2021 to 4/30/2021\n\n<br>\n\n#####  After the train data is updated\n\nAfter the train data was updated, I tested both.\na. whether August of the previous regular season could be predicted well (2019. except for 2020) \nb. whether the last month of the same season could be predicted well.\n The parameters obtained for each were used to make each prediction, and the final submission was weighted  a and b.\n\n- Validation a: \n  - Training : 4/1/2019 to 7/30/2019 \n  - Validation : 8/1/2019 to 8/31/2019\n  - check accuracy and obtaining parameters in this period\n\n- Validation b: \n  - Training : 4/1/2021 to 6/17/2021\n  - Validation : 6/18/2021 to 7/17/2021\n  - check accuracy and obtaining parameters in this period\n\n- Re-training: \n   Re-train on data from 2021/4/1 to 7/17 with parameters obtained in the above period, respectively.\n\n<br>\n\nFinal sub 1: \nPredicted value of the model made with the parameters obtained in verification a × 0.7 +\nPredicted value of the model made with the parameters obtained in verification b × 0.3\n\nFinal sub 2:\nPredicted value of the model created with the parameters obtained in verification a × 0.3 + Predicted value of the model created with the parameters obtained in verification b × 0.7\n\n<br>\n***\n\n##### postscript\nSince my solution does not use lag-target, I believe that it is inferior to models that use lag-target for training and inference if we only look at the results of the first week.\nHowever, I imagine that the accuracy of these models using lag-targets may gradually decrease in future re-runs, and I expect that there will be a few shake-up/down.\n(But first, I need to get the next rerun running correctly...)\n\nThank you for reading this far!",
      "votes": 25
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1467507": "**[update]**\nIn the end, I got 11th place and I was able to stay within the gold medal range. \nI'd like to thank the people running this competition and all the people I competed with! Thank you very much.\n\n\na first week of RE-RUN is over. \nI was lucky enough to have a successful rerun.\nI would like to thank the admin and all the participants for organizing this event.\nThere is still a month to go before it ends, but before I forget, I'll describe my solution.\n(I'm not very good at English, so I basically translate Japanese at DeepL. I apologize if you have trouble reading.)\n<br>\n\n\n***\n\n\n####  Difficult points of this competition\n1. The target is scaled from 0 to 100 on a daily basis.\n1. There is a difference in the information available to pitchers and batters.(The roles of the pitcher and batter are fundamentally different.)\n\n<br>\n\n#### My approach to the above points\n- **Use only in-season data for training**\nWhen I checked the daily median values, I found that there was a considerable difference in the median values between the on-season and off-season.\nI thought that on-season engagement would be determined by performance and success during the season, while off-season engagement would be determined by media exposure and social networking regardless of the season(and we can't use this kind of data).\nIn this competition, we predicts the engagement during the season, so I judged that data from the off-season would be noisy and did not use it for training and validation.\n\n- **Create features that represent what happened on any given day and the game**\nI thought that daily median engagement scores might differ depending on what happened in the game as a whole and what happened in other games on that day.\nTherefore, I created the features related to the actions that occurred throughout the day.\n  - Number of games played that day, total runs scored, home runs, hits, etc.\n  - The average pct of both teams in each game (whether there was an exciting game on the same day between teams with high winning percentages)\n  - The total twitter followers of the teams playing in the game (home and away) as a percentage of the total followers of all teams (whether the game was played between teams that are popular on social networking sites)\n  - The total number of twitter followers of the players who played in the game that day as a percentage of the total number of followers of all players (how influential are the players who played that day on SNS)\n  - How long has the season elapsed (what percentage of the schedule has been completed)?<br>\nIn addition to the above, I added the player's performance.<br><br>\n\n  - Daily performance\n  - Season performance\n  - Previous season performance (last season's final performance)\n\n- **Built three different models**\none for pitchers, one for batters, and one according to no play for the 2021 season.\nThe features used in each model is slightly different from each other. \nThe features for any given day or game described above are the same for all models, \nbut in the pitcher's model, the only feature used in the player's performance was the performance related to pitching. in the batter's model, only the performance related to batting was used. \nFor the model of a player without a season of play, we built the model without using any features related to performance.<br>\nThe reason for this, as mentioned earlier, is that pitchers and batters have distinctly different roles. For example, even if a pitcher gets three hits, his reputation will not be high if he is scored on by his crucial pitches.\nIn addition, for players who did not play in the season, we could not obtain their performance in the current season at all and had no material to judge them, so we built a separate model with fewer features.(no use performance features) \n<br>\nAs a side note, I don't use target-lag-feature.\nWhen I was testing the three models separately, the accuracy of the pitcher and batter models decreased when using past targets, so I decided not to use variables related to past targets this competition.\n\n\n<br>\n\n#### validation scheme\nFor training and validation,  I used data from each season.\nIn addition, I did not use the data for 2020 because the number of games was small and we thought that there might be abnormal movements due to the corona disaster.\nAll models are made with lightgbm.\n\n<br>\n\n##### Before the train data was updated (LB:1.3104(Previous leaderboard))\n - Training : 4/1/2019 to 4/30/2019\n - Validation : 5/1/2019 to 5/30/2019 → check accuracy and obtained parameters in this period\n - Re-training: re-train with parameters obtained in the above period on data from 4/1/2021 to 4/30/2021\n\n<br>\n\n#####  After the train data is updated\n\nAfter the train data was updated, I tested both.\na. whether August of the previous regular season could be predicted well (2019. except for 2020) \nb. whether the last month of the same season could be predicted well.\n The parameters obtained for each were used to make each prediction, and the final submission was weighted  a and b.\n\n- Validation a: \n  - Training : 4/1/2019 to 7/30/2019 \n  - Validation : 8/1/2019 to 8/31/2019\n  - check accuracy and obtaining parameters in this period\n\n- Validation b: \n  - Training : 4/1/2021 to 6/17/2021\n  - Validation : 6/18/2021 to 7/17/2021\n  - check accuracy and obtaining parameters in this period\n\n- Re-training: \n   Re-train on data from 2021/4/1 to 7/17 with parameters obtained in the above period, respectively.\n\n<br>\n\nFinal sub 1: \nPredicted value of the model made with the parameters obtained in verification a × 0.7 +\nPredicted value of the model made with the parameters obtained in verification b × 0.3\n\nFinal sub 2:\nPredicted value of the model created with the parameters obtained in verification a × 0.3 + Predicted value of the model created with the parameters obtained in verification b × 0.7\n\n<br>\n***\n\n##### postscript\nSince my solution does not use lag-target, I believe that it is inferior to models that use lag-target for training and inference if we only look at the results of the first week.\nHowever, I imagine that the accuracy of these models using lag-targets may gradually decrease in future re-runs, and I expect that there will be a few shake-up/down.\n(But first, I need to get the next rerun running correctly...)\n\nThank you for reading this far!"
  }
}