{
  "id": 256559,
  "title": "While waiting for the final results.....What solutions/approaches were used",
  "url": "/competitions/mlb-player-digital-engagement-forecasting/discussion/256559",
  "author_name": "",
  "post_date": "2021-08-02T09:01:01.594199Z",
  "votes": 15,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Heading for the first and ongoing evaluation period, it’s sure interesting to see the results coming.<br>\nParallel to this it would be also interesting to hear short comments about what approaches are used.</p>\n<p>My 2 simple approaches / subs, based on some public codes (credit to those who contributed to the community) but did further tuning and retrained and some code optimizing for the evaluation period:</p>\n<p>1<br>\nSingle LGBM but put in a code that checked if there are newer data in the updated train file, if so, uses that for new feature engineering and training, hence the single lgbm to save time and space. But didn’t used the whole data for post processing only the part that was new to saving time and memory and merged it with the previous data.<br>\nThe LGBM run a training with the latest data and 45 days ahead mae validation, when finished used the best iteration to update the parameters for a new model and training with the full data.</p>\n<p>I also used not only memory garbage cleaning but also a totally refresh of memory, so for this I saved all necessary file from the training and after the reset/refresh I loaded only the data needed to the memory for the inference.</p>\n<p>2<br>\nMore models, lgbm, catboost, NN, for this inference only pretrained models were used, on the latest 0717 version. One of the lgbm was a copy the first sub but pretrained on the full data, the rest only used the validation versions, all with 45 days validation. For the NN I used 30 models, different epochs, some trained with validation following a short final epochs with full data after validation stop, and also models with different data approaches etc.</p>",
  "messages": [
    {
      "id": "1408022",
      "postDate": "08/02/2021 09:01:01",
      "content": "<p>Heading for the first and ongoing evaluation period, it’s sure interesting to see the results coming.<br>\nParallel to this it would be also interesting to hear short comments about what approaches are used.</p>\n<p>My 2 simple approaches / subs, based on some public codes (credit to those who contributed to the community) but did further tuning and retrained and some code optimizing for the evaluation period:</p>\n<p>1<br>\nSingle LGBM but put in a code that checked if there are newer data in the updated train file, if so, uses that for new feature engineering and training, hence the single lgbm to save time and space. But didn’t used the whole data for post processing only the part that was new to saving time and memory and merged it with the previous data.<br>\nThe LGBM run a training with the latest data and 45 days ahead mae validation, when finished used the best iteration to update the parameters for a new model and training with the full data.</p>\n<p>I also used not only memory garbage cleaning but also a totally refresh of memory, so for this I saved all necessary file from the training and after the reset/refresh I loaded only the data needed to the memory for the inference.</p>\n<p>2<br>\nMore models, lgbm, catboost, NN, for this inference only pretrained models were used, on the latest 0717 version. One of the lgbm was a copy the first sub but pretrained on the full data, the rest only used the validation versions, all with 45 days validation. For the NN I used 30 models, different epochs, some trained with validation following a short final epochs with full data after validation stop, and also models with different data approaches etc.</p>",
      "rawMarkdown": "Heading for the first and ongoing evaluation period, it’s sure interesting to see the results coming.\nParallel to this it would be also interesting to hear short comments about what approaches are used.\n\nMy 2 simple approaches / subs, based on some public codes (credit to those who contributed to the community) but did further tuning and retrained and some code optimizing for the evaluation period:\n\n1\nSingle LGBM but put in a code that checked if there are newer data in the updated train file, if so, uses that for new feature engineering and training, hence the single lgbm to save time and space. But didn’t used the whole data for post processing only the part that was new to saving time and memory and merged it with the previous data.\nThe LGBM run a training with the latest data and 45 days ahead mae validation, when finished used the best iteration to update the parameters for a new model and training with the full data.\n\nI also used not only memory garbage cleaning but also a totally refresh of memory, so for this I saved all necessary file from the training and after the reset/refresh I loaded only the data needed to the memory for the inference.\n\n2\nMore models, lgbm, catboost, NN, for this inference only pretrained models were used, on the latest 0717 version. One of the lgbm was a copy the first sub but pretrained on the full data, the rest only used the validation versions, all with 45 days validation. For the NN I used 30 models, different epochs, some trained with validation following a short final epochs with full data after validation stop, and also models with different data approaches etc.",
      "votes": null
    },
    {
      "id": "1408742",
      "postDate": "08/02/2021 15:26:18",
      "content": "<p>I spent majority of my time building efficient feature generation pipeline (Maybe little too much! I hope nothing fails on test run). The end output of all that software engineering is that <strong>I am able to generate ~600 features on train data in 5 minutes</strong> 😁.  Which essentially means, I can run training from scratch on new data updated till July 31st.</p>\n<p>Strategy for 2 submissions were as follows:</p>\n<p>Robust model: An ensemble of 6 LGB model (trained on different time periods and different feature sets) and 8 NN models. The validation score on different time frames is comparable for all models, (every time frame a different model is best)</p>\n<p>Overfit to most recent data: I selected a notebook that will train 2 LGB models on all data till July 31st (till July 17th it took 5 hours to run so I might just squeeze in 6 hours of limit). To make results stable, added results of robust model with 50% weight.</p>\n<p>Didn't spend much time on manual feature engineering, just threw kitchen sink at my models.</p>\n<p>I plan to write detailed post on <strong>efficient feature engineering pipeline for time x user</strong> kind of datasets</p>",
      "rawMarkdown": "I spent majority of my time building efficient feature generation pipeline (Maybe little too much! I hope nothing fails on test run). The end output of all that software engineering is that **I am able to generate ~600 features on train data in 5 minutes** 😁.  Which essentially means, I can run training from scratch on new data updated till July 31st.\n\nStrategy for 2 submissions were as follows:\n\nRobust model: An ensemble of 6 LGB model (trained on different time periods and different feature sets) and 8 NN models. The validation score on different time frames is comparable for all models, (every time frame a different model is best)\n\nOverfit to most recent data: I selected a notebook that will train 2 LGB models on all data till July 31st (till July 17th it took 5 hours to run so I might just squeeze in 6 hours of limit). To make results stable, added results of robust model with 50% weight.\n\nDidn't spend much time on manual feature engineering, just threw kitchen sink at my models.\n\nI plan to write detailed post on **efficient feature engineering pipeline for time x user** kind of datasets",
      "votes": null
    },
    {
      "id": "1408941",
      "postDate": "08/02/2021 16:23:04",
      "content": "<p>Thanks for sharing, sound like a model to follow! Fast post-processing of the new data. Hope the timelimit works for your training with the bigger set. I had an idea to also include random forest as the target curve is a min.max curve over time to minimize to far away prediction, but didn’t come so far in the to-do list ;)</p>",
      "rawMarkdown": "Thanks for sharing, sound like a model to follow! Fast post-processing of the new data. Hope the timelimit works for your training with the bigger set. I had an idea to also include random forest as the target curve is a min.max curve over time to minimize to far away prediction, but didn’t come so far in the to-do list ;)",
      "votes": null
    },
    {
      "id": "1409176",
      "postDate": "08/02/2021 16:40:22",
      "content": "<p>I used one 1D CNN and 2 LGBs. My final prediction is median of these 3 models. LGBs have the same features but one has Huber loss on actual target, the other has l2 loss on log(target). My feature set is very minimal. I pruned a lot of features using LOFO. <a href=\"https://www.kaggle.com/aerdem4/mlb-lofo-feature-importance\" target=\"_blank\">https://www.kaggle.com/aerdem4/mlb-lofo-feature-importance</a></p>\n<p>Local validation: ~1.28<br>\nPublic LB (no Leak): 1.2875</p>",
      "rawMarkdown": "I used one 1D CNN and 2 LGBs. My final prediction is median of these 3 models. LGBs have the same features but one has Huber loss on actual target, the other has l2 loss on log(target). My feature set is very minimal. I pruned a lot of features using LOFO. https://www.kaggle.com/aerdem4/mlb-lofo-feature-importance\n\nLocal validation: ~1.28\nPublic LB (no Leak): 1.2875",
      "votes": null
    },
    {
      "id": "1409598",
      "postDate": "08/02/2021 17:04:33",
      "content": "<p>Btw really nice notebook on LOFO. May I ask what is your validation set?</p>",
      "rawMarkdown": "Btw really nice notebook on LOFO. May I ask what is your validation set?",
      "votes": null
    },
    {
      "id": "1409713",
      "postDate": "08/02/2021 17:11:06",
      "content": "<p>Thank you. I have 5 time splits August 2019, September 2019, August 2020, September 2020, April 2021 being validation sets.</p>",
      "rawMarkdown": "Thank you. I have 5 time splits August 2019, September 2019, August 2020, September 2020, April 2021 being validation sets.",
      "votes": null
    },
    {
      "id": "1409854",
      "postDate": "08/02/2021 17:18:40",
      "content": "<p>Smart time split! And thanks for sharing 👍</p>",
      "rawMarkdown": "Smart time split! And thanks for sharing 👍",
      "votes": null
    },
    {
      "id": "1416537",
      "postDate": "08/02/2021 23:10:14",
      "content": "<p>We focused to improve a single LGB model, one of our final selections (and best on public LB) is an ensemble of multiple LGB models, trained using the same set of features, we relied on feature engineering to improve our score. No lag features were used in our best LGB model, hope our features work in private LB.</p>",
      "rawMarkdown": "We focused to improve a single LGB model, one of our final selections (and best on public LB) is an ensemble of multiple LGB models, trained using the same set of features, we relied on feature engineering to improve our score. No lag features were used in our best LGB model, hope our features work in private LB.",
      "votes": null
    },
    {
      "id": "1439069",
      "postDate": "08/03/2021 16:44:03",
      "content": "<p>May I ask what's your best CV/LB score for a single LGBM model?</p>",
      "rawMarkdown": "May I ask what's your best CV/LB score for a single LGBM model?",
      "votes": null
    },
    {
      "id": "1439104",
      "postDate": "08/03/2021 16:45:18",
      "content": "<p>May I ask by how much percentage you pruned the features?</p>",
      "rawMarkdown": "May I ask by how much percentage you pruned the features?",
      "votes": null
    },
    {
      "id": "1443994",
      "postDate": "08/04/2021 04:44:15",
      "content": "<p>I used \"gamesStartedPitching lag feature\". Starting pitcher sometimes pitch (one time per 4~5days) but they get high target1~4. I made sub model of predicting \"this pitcher will pitch tomorrow\". If this pitcher will pitch tomorrow, he will get high target.<br>\n\"This pitcher will pitch tomorrow\" is easy to predict, because MLB pitchers have rotation (almost of them pitch one time per 4~5days)</p>",
      "rawMarkdown": "I used \"gamesStartedPitching lag feature\". Starting pitcher sometimes pitch (one time per 4~5days) but they get high target1~4. I made sub model of predicting \"this pitcher will pitch tomorrow\". If this pitcher will pitch tomorrow, he will get high target.\n\"This pitcher will pitch tomorrow\" is easy to predict, because MLB pitchers have rotation (almost of them pitch one time per 4~5days)",
      "votes": null
    },
    {
      "id": "1444830",
      "postDate": "08/04/2021 07:05:40",
      "content": "<p>Thanks for sharing! Nice real-world-insight-feature!</p>",
      "rawMarkdown": "Thanks for sharing! Nice real-world-insight-feature!",
      "votes": null
    },
    {
      "id": "1446203",
      "postDate": "08/04/2021 11:26:08",
      "content": "<p>Hi did you use an expanding window for your split?</p>",
      "rawMarkdown": "Hi did you use an expanding window for your split?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1408742,
      "author_name": "tezdhar",
      "author_url": "",
      "post_date": "08/02/2021 15:26:18",
      "content": "<p>I spent majority of my time building efficient feature generation pipeline (Maybe little too much! I hope nothing fails on test run). The end output of all that software engineering is that <strong>I am able to generate ~600 features on train data in 5 minutes</strong> 😁.  Which essentially means, I can run training from scratch on new data updated till July 31st.</p>\n<p>Strategy for 2 submissions were as follows:</p>\n<p>Robust model: An ensemble of 6 LGB model (trained on different time periods and different feature sets) and 8 NN models. The validation score on different time frames is comparable for all models, (every time frame a different model is best)</p>\n<p>Overfit to most recent data: I selected a notebook that will train 2 LGB models on all data till July 31st (till July 17th it took 5 hours to run so I might just squeeze in 6 hours of limit). To make results stable, added results of robust model with 50% weight.</p>\n<p>Didn't spend much time on manual feature engineering, just threw kitchen sink at my models.</p>\n<p>I plan to write detailed post on <strong>efficient feature engineering pipeline for time x user</strong> kind of datasets</p>",
      "votes": null,
      "replies": [
        {
          "id": 1408941,
          "author_name": "kirderf",
          "author_url": "",
          "post_date": "08/02/2021 16:23:04",
          "content": "<p>Thanks for sharing, sound like a model to follow! Fast post-processing of the new data. Hope the timelimit works for your training with the bigger set. I had an idea to also include random forest as the target curve is a min.max curve over time to minimize to far away prediction, but didn’t come so far in the to-do list ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1409176,
      "author_name": "aerdem4",
      "author_url": "",
      "post_date": "08/02/2021 16:40:22",
      "content": "<p>I used one 1D CNN and 2 LGBs. My final prediction is median of these 3 models. LGBs have the same features but one has Huber loss on actual target, the other has l2 loss on log(target). My feature set is very minimal. I pruned a lot of features using LOFO. <a href=\"https://www.kaggle.com/aerdem4/mlb-lofo-feature-importance\" target=\"_blank\">https://www.kaggle.com/aerdem4/mlb-lofo-feature-importance</a></p>\n<p>Local validation: ~1.28<br>\nPublic LB (no Leak): 1.2875</p>",
      "votes": null,
      "replies": [
        {
          "id": 1409598,
          "author_name": "tezdhar",
          "author_url": "",
          "post_date": "08/02/2021 17:04:33",
          "content": "<p>Btw really nice notebook on LOFO. May I ask what is your validation set?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1409713,
          "author_name": "aerdem4",
          "author_url": "",
          "post_date": "08/02/2021 17:11:06",
          "content": "<p>Thank you. I have 5 time splits August 2019, September 2019, August 2020, September 2020, April 2021 being validation sets.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1409854,
          "author_name": "kirderf",
          "author_url": "",
          "post_date": "08/02/2021 17:18:40",
          "content": "<p>Smart time split! And thanks for sharing 👍</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1439104,
          "author_name": "zacchaeus",
          "author_url": "",
          "post_date": "08/03/2021 16:45:18",
          "content": "<p>May I ask by how much percentage you pruned the features?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1446203,
          "author_name": "toxicmaze",
          "author_url": "",
          "post_date": "08/04/2021 11:26:08",
          "content": "<p>Hi did you use an expanding window for your split?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1416537,
      "author_name": "chinta",
      "author_url": "",
      "post_date": "08/02/2021 23:10:14",
      "content": "<p>We focused to improve a single LGB model, one of our final selections (and best on public LB) is an ensemble of multiple LGB models, trained using the same set of features, we relied on feature engineering to improve our score. No lag features were used in our best LGB model, hope our features work in private LB.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1439069,
          "author_name": "zacchaeus",
          "author_url": "",
          "post_date": "08/03/2021 16:44:03",
          "content": "<p>May I ask what's your best CV/LB score for a single LGBM model?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1443994,
      "author_name": "deepkun1995",
      "author_url": "",
      "post_date": "08/04/2021 04:44:15",
      "content": "<p>I used \"gamesStartedPitching lag feature\". Starting pitcher sometimes pitch (one time per 4~5days) but they get high target1~4. I made sub model of predicting \"this pitcher will pitch tomorrow\". If this pitcher will pitch tomorrow, he will get high target.<br>\n\"This pitcher will pitch tomorrow\" is easy to predict, because MLB pitchers have rotation (almost of them pitch one time per 4~5days)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1444830,
          "author_name": "kirderf",
          "author_url": "",
          "post_date": "08/04/2021 07:05:40",
          "content": "<p>Thanks for sharing! Nice real-world-insight-feature!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1408022": "Heading for the first and ongoing evaluation period, it’s sure interesting to see the results coming.\nParallel to this it would be also interesting to hear short comments about what approaches are used.\n\nMy 2 simple approaches / subs, based on some public codes (credit to those who contributed to the community) but did further tuning and retrained and some code optimizing for the evaluation period:\n\n1\nSingle LGBM but put in a code that checked if there are newer data in the updated train file, if so, uses that for new feature engineering and training, hence the single lgbm to save time and space. But didn’t used the whole data for post processing only the part that was new to saving time and memory and merged it with the previous data.\nThe LGBM run a training with the latest data and 45 days ahead mae validation, when finished used the best iteration to update the parameters for a new model and training with the full data.\n\nI also used not only memory garbage cleaning but also a totally refresh of memory, so for this I saved all necessary file from the training and after the reset/refresh I loaded only the data needed to the memory for the inference.\n\n2\nMore models, lgbm, catboost, NN, for this inference only pretrained models were used, on the latest 0717 version. One of the lgbm was a copy the first sub but pretrained on the full data, the rest only used the validation versions, all with 45 days validation. For the NN I used 30 models, different epochs, some trained with validation following a short final epochs with full data after validation stop, and also models with different data approaches etc.",
    "1408742": "I spent majority of my time building efficient feature generation pipeline (Maybe little too much! I hope nothing fails on test run). The end output of all that software engineering is that **I am able to generate ~600 features on train data in 5 minutes** 😁.  Which essentially means, I can run training from scratch on new data updated till July 31st.\n\nStrategy for 2 submissions were as follows:\n\nRobust model: An ensemble of 6 LGB model (trained on different time periods and different feature sets) and 8 NN models. The validation score on different time frames is comparable for all models, (every time frame a different model is best)\n\nOverfit to most recent data: I selected a notebook that will train 2 LGB models on all data till July 31st (till July 17th it took 5 hours to run so I might just squeeze in 6 hours of limit). To make results stable, added results of robust model with 50% weight.\n\nDidn't spend much time on manual feature engineering, just threw kitchen sink at my models.\n\nI plan to write detailed post on **efficient feature engineering pipeline for time x user** kind of datasets",
    "1408941": "Thanks for sharing, sound like a model to follow! Fast post-processing of the new data. Hope the timelimit works for your training with the bigger set. I had an idea to also include random forest as the target curve is a min.max curve over time to minimize to far away prediction, but didn’t come so far in the to-do list ;)",
    "1409176": "I used one 1D CNN and 2 LGBs. My final prediction is median of these 3 models. LGBs have the same features but one has Huber loss on actual target, the other has l2 loss on log(target). My feature set is very minimal. I pruned a lot of features using LOFO. https://www.kaggle.com/aerdem4/mlb-lofo-feature-importance\n\nLocal validation: ~1.28\nPublic LB (no Leak): 1.2875",
    "1409598": "Btw really nice notebook on LOFO. May I ask what is your validation set?",
    "1409713": "Thank you. I have 5 time splits August 2019, September 2019, August 2020, September 2020, April 2021 being validation sets.",
    "1409854": "Smart time split! And thanks for sharing 👍",
    "1416537": "We focused to improve a single LGB model, one of our final selections (and best on public LB) is an ensemble of multiple LGB models, trained using the same set of features, we relied on feature engineering to improve our score. No lag features were used in our best LGB model, hope our features work in private LB.",
    "1439069": "May I ask what's your best CV/LB score for a single LGBM model?",
    "1439104": "May I ask by how much percentage you pruned the features?",
    "1443994": "I used \"gamesStartedPitching lag feature\". Starting pitcher sometimes pitch (one time per 4~5days) but they get high target1~4. I made sub model of predicting \"this pitcher will pitch tomorrow\". If this pitcher will pitch tomorrow, he will get high target.\n\"This pitcher will pitch tomorrow\" is easy to predict, because MLB pitchers have rotation (almost of them pitch one time per 4~5days)",
    "1444830": "Thanks for sharing! Nice real-world-insight-feature!",
    "1446203": "Hi did you use an expanding window for your split?"
  },
  "source": "meta"
}