{
  "id": 261020,
  "title": "22th solution",
  "url": "/competitions/mlb-player-digital-engagement-forecasting/writeups/chumajin-w-o-leak-1-3019-22th-solution",
  "author_name": "",
  "post_date": "2021-09-10T11:35:22.603Z",
  "votes": 15,
  "comment_count": 2,
  "views": 0,
  "content": "<p>First of all, thank you very much to those who supported me, upvoted EDA, and excited me! I enjoyed experiencing various things.　The top ones are really amazing. I respect them.</p>\n<p>I shared my solution in detail. (日本語解説ありです). <br>\nThis is the 12th solution before the train data is updated(leaked).</p>\n<p><a href=\"https://www.kaggle.com/chumajin/12th-before-train-update-solution-english\" target=\"_blank\">https://www.kaggle.com/chumajin/12th-before-train-update-solution-english</a></p>\n<p>But I don't know the actual result. If I got worse score or submission error, please laugh me… ( I will laugh at my own)</p>\n<p>[Short summary]</p>\n<ul>\n<li><p>I used only optuna and LGBM. It was Public LB1.3019 by ensemble what was created by changing the features.</p></li>\n<li><p>CV is 5 kfold by the average of from target1 to 4.</p></li>\n<li><p>In my 1st phase, I used the merged code that was published by kaggle staff. This LGBM model got 1.3490 score.</p></li>\n<li><p>In my 2nd phase, I added the features about the statics of target values more than 31 days ago because we know the correct answer. This model got better to 1.3373. I used GCP because of memory insufficient.</p></li>\n<li><p>In my 3rd phase, I used the log scale of target value and omitted 0 and 100 value. I fount the clean histogram if I use the log scale of target values. This model got better to 1.3256.</p></li>\n<li><p>I ensemble models made with other features and models made for each position.<br>\nThis models got better to 1.3144.</p></li>\n<li><p>Moreover, no hitter is very high targets value. So I correct it(1.3073). And I found Shohei Ohtani found that LGBM's predictions did not match, and corrected the difference(maybe this is overfit).<br>\nFinally I got to 1.3019.</p></li>\n</ul>\n<p>Thank you so much, good luck for everyone !</p>",
  "messages": [
    {
      "id": "1443258",
      "postDate": "08/04/2021 02:20:55",
      "content": "<p>First of all, thank you very much to those who supported me, upvoted EDA, and excited me! I enjoyed experiencing various things.　The top ones are really amazing. I respect them.</p>\n<p>I shared my solution in detail. (日本語解説ありです). <br>\nThis is the 12th solution before the train data is updated(leaked).</p>\n<p><a href=\"https://www.kaggle.com/chumajin/12th-before-train-update-solution-english\" target=\"_blank\">https://www.kaggle.com/chumajin/12th-before-train-update-solution-english</a></p>\n<p>But I don't know the actual result. If I got worse score or submission error, please laugh me… ( I will laugh at my own)</p>\n<p>[Short summary]</p>\n<ul>\n<li><p>I used only optuna and LGBM. It was Public LB1.3019 by ensemble what was created by changing the features.</p></li>\n<li><p>CV is 5 kfold by the average of from target1 to 4.</p></li>\n<li><p>In my 1st phase, I used the merged code that was published by kaggle staff. This LGBM model got 1.3490 score.</p></li>\n<li><p>In my 2nd phase, I added the features about the statics of target values more than 31 days ago because we know the correct answer. This model got better to 1.3373. I used GCP because of memory insufficient.</p></li>\n<li><p>In my 3rd phase, I used the log scale of target value and omitted 0 and 100 value. I fount the clean histogram if I use the log scale of target values. This model got better to 1.3256.</p></li>\n<li><p>I ensemble models made with other features and models made for each position.<br>\nThis models got better to 1.3144.</p></li>\n<li><p>Moreover, no hitter is very high targets value. So I correct it(1.3073). And I found Shohei Ohtani found that LGBM's predictions did not match, and corrected the difference(maybe this is overfit).<br>\nFinally I got to 1.3019.</p></li>\n</ul>\n<p>Thank you so much, good luck for everyone !</p>",
      "rawMarkdown": "First of all, thank you very much to those who supported me, upvoted EDA, and excited me! I enjoyed experiencing various things.　The top ones are really amazing. I respect them.\n\nI shared my solution in detail. (日本語解説ありです). \nThis is the 12th solution before the train data is updated(leaked).\n\nhttps://www.kaggle.com/chumajin/12th-before-train-update-solution-english\n\n\nBut I don't know the actual result. If I got worse score or submission error, please laugh me... ( I will laugh at my own)\n\n\n[Short summary]\n* I used only optuna and LGBM. It was Public LB1.3019 by ensemble what was created by changing the features.\n\n* CV is 5 kfold by the average of from target1 to 4.\n\n* In my 1st phase, I used the merged code that was published by kaggle staff. This LGBM model got 1.3490 score.\n\n* In my 2nd phase, I added the features about the statics of target values more than 31 days ago because we know the correct answer. This model got better to 1.3373. I used GCP because of memory insufficient.\n\n* In my 3rd phase, I used the log scale of target value and omitted 0 and 100 value. I fount the clean histogram if I use the log scale of target values. This model got better to 1.3256.\n\n*  I ensemble models made with other features and models made for each position.\nThis models got better to 1.3144.\n\n* Moreover, no hitter is very high targets value. So I correct it(1.3073). And I found Shohei Ohtani found that LGBM's predictions did not match, and corrected the difference(maybe this is overfit).\nFinally I got to 1.3019.\n\nThank you so much, good luck for everyone !",
      "votes": null
    },
    {
      "id": "1443453",
      "postDate": "08/04/2021 03:03:30",
      "content": "<p>Thanks for sharing. Can you elaborate the last bullet point?</p>",
      "rawMarkdown": "Thanks for sharing. Can you elaborate the last bullet point?",
      "votes": null
    },
    {
      "id": "1445360",
      "postDate": "08/04/2021 08:38:23",
      "content": "<p>Thank you for comments.</p>\n<p>This is the analyzed result of no hitter.</p>\n<p><img src=\"https://raw.githubusercontent.com/chumajin/test/main/nohitter.JPG\" alt=\"\"></p>\n<p>if no hitter is 1, target is very high. So, I adopted this median values when no hitter is 1.</p>\n<p>This is the validation result of 2021 April. The difference of predictions and grand truth for Shohei Ohtani was too big. So I add the difference of it to prediction. But this is maybe not good because the score was getting better only in the case of Shohei Ohtani. Honestly, I tried to making models for Shohei Ohtani, but I could not.</p>\n<p><img src=\"https://raw.githubusercontent.com/chumajin/test/main/Shohei.JPG\" alt=\"\"> </p>",
      "rawMarkdown": "Thank you for comments.\n\nThis is the analyzed result of no hitter.\n\n![](https://raw.githubusercontent.com/chumajin/test/main/nohitter.JPG)\n\nif no hitter is 1, target is very high. So, I adopted this median values when no hitter is 1.\n\nThis is the validation result of 2021 April. The difference of predictions and grand truth for Shohei Ohtani was too big. So I add the difference of it to prediction. But this is maybe not good because the score was getting better only in the case of Shohei Ohtani. Honestly, I tried to making models for Shohei Ohtani, but I could not.\n\n![](https://raw.githubusercontent.com/chumajin/test/main/Shohei.JPG)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1443453,
      "author_name": "zacchaeus",
      "author_url": "",
      "post_date": "08/04/2021 03:03:30",
      "content": "<p>Thanks for sharing. Can you elaborate the last bullet point?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1445360,
          "author_name": "chumajin",
          "author_url": "",
          "post_date": "08/04/2021 08:38:23",
          "content": "<p>Thank you for comments.</p>\n<p>This is the analyzed result of no hitter.</p>\n<p><img src=\"https://raw.githubusercontent.com/chumajin/test/main/nohitter.JPG\" alt=\"\"></p>\n<p>if no hitter is 1, target is very high. So, I adopted this median values when no hitter is 1.</p>\n<p>This is the validation result of 2021 April. The difference of predictions and grand truth for Shohei Ohtani was too big. So I add the difference of it to prediction. But this is maybe not good because the score was getting better only in the case of Shohei Ohtani. Honestly, I tried to making models for Shohei Ohtani, but I could not.</p>\n<p><img src=\"https://raw.githubusercontent.com/chumajin/test/main/Shohei.JPG\" alt=\"\"> </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1443258": "First of all, thank you very much to those who supported me, upvoted EDA, and excited me! I enjoyed experiencing various things.　The top ones are really amazing. I respect them.\n\nI shared my solution in detail. (日本語解説ありです). \nThis is the 12th solution before the train data is updated(leaked).\n\nhttps://www.kaggle.com/chumajin/12th-before-train-update-solution-english\n\n\nBut I don't know the actual result. If I got worse score or submission error, please laugh me... ( I will laugh at my own)\n\n\n[Short summary]\n* I used only optuna and LGBM. It was Public LB1.3019 by ensemble what was created by changing the features.\n\n* CV is 5 kfold by the average of from target1 to 4.\n\n* In my 1st phase, I used the merged code that was published by kaggle staff. This LGBM model got 1.3490 score.\n\n* In my 2nd phase, I added the features about the statics of target values more than 31 days ago because we know the correct answer. This model got better to 1.3373. I used GCP because of memory insufficient.\n\n* In my 3rd phase, I used the log scale of target value and omitted 0 and 100 value. I fount the clean histogram if I use the log scale of target values. This model got better to 1.3256.\n\n*  I ensemble models made with other features and models made for each position.\nThis models got better to 1.3144.\n\n* Moreover, no hitter is very high targets value. So I correct it(1.3073). And I found Shohei Ohtani found that LGBM's predictions did not match, and corrected the difference(maybe this is overfit).\nFinally I got to 1.3019.\n\nThank you so much, good luck for everyone !",
    "1443453": "Thanks for sharing. Can you elaborate the last bullet point?",
    "1445360": "Thank you for comments.\n\nThis is the analyzed result of no hitter.\n\n![](https://raw.githubusercontent.com/chumajin/test/main/nohitter.JPG)\n\nif no hitter is 1, target is very high. So, I adopted this median values when no hitter is 1.\n\nThis is the validation result of 2021 April. The difference of predictions and grand truth for Shohei Ohtani was too big. So I add the difference of it to prediction. But this is maybe not good because the score was getting better only in the case of Shohei Ohtani. Honestly, I tried to making models for Shohei Ohtani, but I could not.\n\n![](https://raw.githubusercontent.com/chumajin/test/main/Shohei.JPG)"
  },
  "source": "meta"
}