{
  "id": 271345,
  "title": "5th Place Solution : My First Competition and First Gold ",
  "url": "/competitions/mlb-player-digital-engagement-forecasting/writeups/blife-mariko-5th-place-solution-my-first-competiti",
  "author_name": "",
  "post_date": "2021-09-12T22:16:49.430Z",
  "votes": 24,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi all</p>\n<p>I would like to thank kaggle and the organizers for such a good competition. <br>\nI also thank  my teammate(<a href=\"https://www.kaggle.com/shinnyayoshida\" target=\"_blank\">@Hyper-Positive-Yancy</a>), he did some EDA and tuned NN.</p>\n<h1>Models Used</h1>\n<p>Our final ensemble consisted of</p>\n<ul>\n<li>Lightgbm X 8</li>\n<li>CATBOOST X 4</li>\n<li>ANN X1</li>\n</ul>\n<h1>Data usage period</h1>\n<ul>\n<li><strong>in-season sampling</strong><br>\nI only use in-season data.<br>\nEven if I extracted out-of-season data from our data, we could not confirm any deterioration in accuracy.<br>\nBut, In 2019 data , I eliminated a lot of data. In this year, The retirement match of the great Ichiro Suzuki was held in Japan. He didnt play well in the retirement game, but he had a high engagement.<br>\nI considered this data outlier and deleted it. From this fact, I Concluded, Special matches(like retirement match) should be removed from the data</li>\n</ul>\n<h1>Feature engineering</h1>\n<ul>\n<li><p><strong>gamesStartedPitching lag feature</strong><br>\nAs I mentioned <a href=\"https://www.kaggle.com/c/mlb-player-digital-engagement-forecasting/discussion/261357#1446144\" target=\"_blank\">here</a><br>\nI used \"gamesStartedPitching lag feature\". Starting pitcher sometimes pitch (one time per 4~5days) but they get high target1~4. I made sub model of predicting \"this pitcher will pitch tomorrow\". If this pitcher will pitch tomorrow, he will get high target.<br>\n\"This pitcher will pitch tomorrow\" is easy to predict, because MLB pitchers have rotation (almost of them pitch one time per 4~5days)<br>\nwhen I use this feature with \"target-lag-feature\", my public LB score became better. </p></li>\n<li><p><strong>tomorrow game exists or not</strong><br>\nIn this competition, we predict \"tomorrow\" engagement, but stats of player is today stats. Even if the players did well today, we could see that the engagement value would be low if there was no match tomorrow.<br>\nSo, Whether or not there will be a match tomorrow will be an important factor in predicting engagement</p></li>\n<li><p><strong>target-lag-feature</strong><br>\nAs I mentioned before, this feature is effective when used with \"gamesStartedPitching lag feature\". Therefore, I thought it would be very meaningful to use the target-lag-feature of  3 to 7 days.<br>\nBut, Second half of the evaluation period, I have to use predicted value of target data.<br>\nIt may be one of the factors that worsen our second half model performance.</p></li>\n<li><p><strong>batter contribution</strong></p></li>\n<li><p><strong>pitching contribution</strong><br>\nIn order to evaluate all athletes fairly, we used the evaluation index for athletes.<br>\nI mainly use two evaluation index \" batter contribution\" and \"pitching contribution\"<br>\nI found  \" batter contribution\" <a href=\"http://maddog31.xyz/baseball-web/contribution_degree/contribution_degree2/\" target=\"_blank\">here</a> sorry this page is written in japanese.<br>\nI found \"pitching game score\"  <a href=\"https://www.mlb.com/glossary/advanced-stats/game-score\" target=\"_blank\">here</a><br>\nAs a problem, even with the same position pitcher, there are types of relief and starting pitcher, and the pitching game score of the relief tends to be low. So the probability that the player will throw as starting pitcher has also been added as a feature.</p></li>\n</ul>\n<h1>validation scheme</h1>\n<p>For training and validation, I used data from each in-season.<br>\nAfter the train data was updated, I changed validation scheme.</p>\n<p>Training : 4/1/2018 to 6/30/2021<br>\nValidation : 7/1/2021 to 7/31/2021</p>",
  "messages": [
    {
      "id": "1508190",
      "postDate": "09/10/2021 01:41:21",
      "content": "<p>Hi all</p>\n<p>I would like to thank kaggle and the organizers for such a good competition. <br>\nI also thank  my teammate(<a href=\"https://www.kaggle.com/shinnyayoshida\" target=\"_blank\">@Hyper-Positive-Yancy</a>), he did some EDA and tuned NN.</p>\n<h1>Models Used</h1>\n<p>Our final ensemble consisted of</p>\n<ul>\n<li>Lightgbm X 8</li>\n<li>CATBOOST X 4</li>\n<li>ANN X1</li>\n</ul>\n<h1>Data usage period</h1>\n<ul>\n<li><strong>in-season sampling</strong><br>\nI only use in-season data.<br>\nEven if I extracted out-of-season data from our data, we could not confirm any deterioration in accuracy.<br>\nBut, In 2019 data , I eliminated a lot of data. In this year, The retirement match of the great Ichiro Suzuki was held in Japan. He didnt play well in the retirement game, but he had a high engagement.<br>\nI considered this data outlier and deleted it. From this fact, I Concluded, Special matches(like retirement match) should be removed from the data</li>\n</ul>\n<h1>Feature engineering</h1>\n<ul>\n<li><p><strong>gamesStartedPitching lag feature</strong><br>\nAs I mentioned <a href=\"https://www.kaggle.com/c/mlb-player-digital-engagement-forecasting/discussion/261357#1446144\" target=\"_blank\">here</a><br>\nI used \"gamesStartedPitching lag feature\". Starting pitcher sometimes pitch (one time per 4~5days) but they get high target1~4. I made sub model of predicting \"this pitcher will pitch tomorrow\". If this pitcher will pitch tomorrow, he will get high target.<br>\n\"This pitcher will pitch tomorrow\" is easy to predict, because MLB pitchers have rotation (almost of them pitch one time per 4~5days)<br>\nwhen I use this feature with \"target-lag-feature\", my public LB score became better. </p></li>\n<li><p><strong>tomorrow game exists or not</strong><br>\nIn this competition, we predict \"tomorrow\" engagement, but stats of player is today stats. Even if the players did well today, we could see that the engagement value would be low if there was no match tomorrow.<br>\nSo, Whether or not there will be a match tomorrow will be an important factor in predicting engagement</p></li>\n<li><p><strong>target-lag-feature</strong><br>\nAs I mentioned before, this feature is effective when used with \"gamesStartedPitching lag feature\". Therefore, I thought it would be very meaningful to use the target-lag-feature of  3 to 7 days.<br>\nBut, Second half of the evaluation period, I have to use predicted value of target data.<br>\nIt may be one of the factors that worsen our second half model performance.</p></li>\n<li><p><strong>batter contribution</strong></p></li>\n<li><p><strong>pitching contribution</strong><br>\nIn order to evaluate all athletes fairly, we used the evaluation index for athletes.<br>\nI mainly use two evaluation index \" batter contribution\" and \"pitching contribution\"<br>\nI found  \" batter contribution\" <a href=\"http://maddog31.xyz/baseball-web/contribution_degree/contribution_degree2/\" target=\"_blank\">here</a> sorry this page is written in japanese.<br>\nI found \"pitching game score\"  <a href=\"https://www.mlb.com/glossary/advanced-stats/game-score\" target=\"_blank\">here</a><br>\nAs a problem, even with the same position pitcher, there are types of relief and starting pitcher, and the pitching game score of the relief tends to be low. So the probability that the player will throw as starting pitcher has also been added as a feature.</p></li>\n</ul>\n<h1>validation scheme</h1>\n<p>For training and validation, I used data from each in-season.<br>\nAfter the train data was updated, I changed validation scheme.</p>\n<p>Training : 4/1/2018 to 6/30/2021<br>\nValidation : 7/1/2021 to 7/31/2021</p>",
      "rawMarkdown": "Hi all\n\nI would like to thank kaggle and the organizers for such a good competition. \nI also thank  my teammate([@Hyper-Positive-Yancy](https://www.kaggle.com/shinnyayoshida)), he did some EDA and tuned NN.\n\n\n# Models Used \nOur final ensemble consisted of\n- Lightgbm X 8\n- CATBOOST X 4\n- ANN X1\n\n# Data usage period\n- **in-season sampling**\nI only use in-season data.\nEven if I extracted out-of-season data from our data, we could not confirm any deterioration in accuracy.\nBut, In 2019 data , I eliminated a lot of data. In this year, The retirement match of the great Ichiro Suzuki was held in Japan. He didnt play well in the retirement game, but he had a high engagement.\nI considered this data outlier and deleted it. From this fact, I Concluded, Special matches(like retirement match) should be removed from the data\n\n# Feature engineering\n- **gamesStartedPitching lag feature**\nAs I mentioned [here](https://www.kaggle.com/c/mlb-player-digital-engagement-forecasting/discussion/261357#1446144)\nI used \"gamesStartedPitching lag feature\". Starting pitcher sometimes pitch (one time per 4~5days) but they get high target1~4. I made sub model of predicting \"this pitcher will pitch tomorrow\". If this pitcher will pitch tomorrow, he will get high target.\n\"This pitcher will pitch tomorrow\" is easy to predict, because MLB pitchers have rotation (almost of them pitch one time per 4~5days)\nwhen I use this feature with \"target-lag-feature\", my public LB score became better. \n\n- **tomorrow game exists or not**\nIn this competition, we predict \"tomorrow\" engagement, but stats of player is today stats. Even if the players did well today, we could see that the engagement value would be low if there was no match tomorrow.\nSo, Whether or not there will be a match tomorrow will be an important factor in predicting engagement\n\n- **target-lag-feature**\nAs I mentioned before, this feature is effective when used with \"gamesStartedPitching lag feature\". Therefore, I thought it would be very meaningful to use the target-lag-feature of  3 to 7 days.\nBut, Second half of the evaluation period, I have to use predicted value of target data.\nIt may be one of the factors that worsen our second half model performance.\n\n- **batter contribution**\n- **pitching contribution**\nIn order to evaluate all athletes fairly, we used the evaluation index for athletes.\nI mainly use two evaluation index \" batter contribution\" and \"pitching contribution\"\nI found  \" batter contribution\" [here] (http://maddog31.xyz/baseball-web/contribution_degree/contribution_degree2/) sorry this page is written in japanese.\nI found \"pitching game score\"  [here](https://www.mlb.com/glossary/advanced-stats/game-score)\nAs a problem, even with the same position pitcher, there are types of relief and starting pitcher, and the pitching game score of the relief tends to be low. So the probability that the player will throw as starting pitcher has also been added as a feature.\n\n\n\n# validation scheme\nFor training and validation, I used data from each in-season.\nAfter the train data was updated, I changed validation scheme.\n\nTraining : 4/1/2018 to 6/30/2021\nValidation : 7/1/2021 to 7/31/2021",
      "votes": null
    },
    {
      "id": "1508303",
      "postDate": "09/10/2021 05:47:23",
      "content": "<p>congrats!👍</p>",
      "rawMarkdown": "congrats!👍",
      "votes": null
    },
    {
      "id": "1508554",
      "postDate": "09/10/2021 10:40:39",
      "content": "<p>thanks<br>\n<a href=\"https://www.kaggle.com/zacchaeus\" target=\"_blank\">@zacchaeus</a> </p>",
      "rawMarkdown": "thanks\n@zacchaeus",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1508303,
      "author_name": "zacchaeus",
      "author_url": "",
      "post_date": "09/10/2021 05:47:23",
      "content": "<p>congrats!👍</p>",
      "votes": null,
      "replies": [
        {
          "id": 1508554,
          "author_name": "deepkun1995",
          "author_url": "",
          "post_date": "09/10/2021 10:40:39",
          "content": "<p>thanks<br>\n<a href=\"https://www.kaggle.com/zacchaeus\" target=\"_blank\">@zacchaeus</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1508190": "Hi all\n\nI would like to thank kaggle and the organizers for such a good competition. \nI also thank  my teammate([@Hyper-Positive-Yancy](https://www.kaggle.com/shinnyayoshida)), he did some EDA and tuned NN.\n\n\n# Models Used \nOur final ensemble consisted of\n- Lightgbm X 8\n- CATBOOST X 4\n- ANN X1\n\n# Data usage period\n- **in-season sampling**\nI only use in-season data.\nEven if I extracted out-of-season data from our data, we could not confirm any deterioration in accuracy.\nBut, In 2019 data , I eliminated a lot of data. In this year, The retirement match of the great Ichiro Suzuki was held in Japan. He didnt play well in the retirement game, but he had a high engagement.\nI considered this data outlier and deleted it. From this fact, I Concluded, Special matches(like retirement match) should be removed from the data\n\n# Feature engineering\n- **gamesStartedPitching lag feature**\nAs I mentioned [here](https://www.kaggle.com/c/mlb-player-digital-engagement-forecasting/discussion/261357#1446144)\nI used \"gamesStartedPitching lag feature\". Starting pitcher sometimes pitch (one time per 4~5days) but they get high target1~4. I made sub model of predicting \"this pitcher will pitch tomorrow\". If this pitcher will pitch tomorrow, he will get high target.\n\"This pitcher will pitch tomorrow\" is easy to predict, because MLB pitchers have rotation (almost of them pitch one time per 4~5days)\nwhen I use this feature with \"target-lag-feature\", my public LB score became better. \n\n- **tomorrow game exists or not**\nIn this competition, we predict \"tomorrow\" engagement, but stats of player is today stats. Even if the players did well today, we could see that the engagement value would be low if there was no match tomorrow.\nSo, Whether or not there will be a match tomorrow will be an important factor in predicting engagement\n\n- **target-lag-feature**\nAs I mentioned before, this feature is effective when used with \"gamesStartedPitching lag feature\". Therefore, I thought it would be very meaningful to use the target-lag-feature of  3 to 7 days.\nBut, Second half of the evaluation period, I have to use predicted value of target data.\nIt may be one of the factors that worsen our second half model performance.\n\n- **batter contribution**\n- **pitching contribution**\nIn order to evaluate all athletes fairly, we used the evaluation index for athletes.\nI mainly use two evaluation index \" batter contribution\" and \"pitching contribution\"\nI found  \" batter contribution\" [here] (http://maddog31.xyz/baseball-web/contribution_degree/contribution_degree2/) sorry this page is written in japanese.\nI found \"pitching game score\"  [here](https://www.mlb.com/glossary/advanced-stats/game-score)\nAs a problem, even with the same position pitcher, there are types of relief and starting pitcher, and the pitching game score of the relief tends to be low. So the probability that the player will throw as starting pitcher has also been added as a feature.\n\n\n\n# validation scheme\nFor training and validation, I used data from each in-season.\nAfter the train data was updated, I changed validation scheme.\n\nTraining : 4/1/2018 to 6/30/2021\nValidation : 7/1/2021 to 7/31/2021",
    "1508303": "congrats!👍",
    "1508554": "thanks\n@zacchaeus"
  },
  "source": "meta"
}