{
  "id": 271683,
  "title": "8th place solution (Moro & taksai)",
  "url": "/competitions/mlb-player-digital-engagement-forecasting/writeups/nakano-8th-place-solution-moro-taksai",
  "author_name": "",
  "post_date": "2021-09-12T03:40:03.800004800Z",
  "votes": 18,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Thank you to the organizers for the fun competition and everyone who participated.<br>\nAnd thank you to my teammate ( taksai <a href=\"https://www.kaggle.com/tsaito21219\" target=\"_blank\">@tsaito21219</a> ).</p>\n<p>I share our team's solution.</p>\n<h1>summary:</h1>\n<ul>\n<li>LSTM model</li>\n<li>valid: 2021-05, 2021-06, 2021-07</li>\n<li>ensemble LSTM:MLP:LightGBM=5:1:1</li>\n</ul>\n<h1>1. validation</h1>\n<ul>\n<li>private-data(2021-08) is in-season, so we use in-season data as validation, and select the month near 2021-08.</li>\n<li>valid: 2021-05, 2021-06, 2021-07</li>\n<li>prepared some patterns as training-data while shifting the period. </li>\n<li>trained 10-fold for each model</li>\n</ul>\n<p><a href=\"https://postimg.cc/nCds5N1C\" target=\"_blank\"><img src=\"https://i.postimg.cc/pLvD9HhY/img1-validation.png\" alt=\"img1-validation.png\"></a></p>\n<h1>2. preprocess data</h1>\n<ul>\n<li>features: total number of 232<ul>\n<li>mean/median/std/min/max of target per playerId last month</li>\n<li>join each table by key(date/playerId/teamId), use the almost feature of each table</li>\n<li>use all tables except events.csv</li>\n<li>not use target-lag-feature</li></ul></li>\n<li>didn't use target-lag-feature, because it's risky. However, if spliting the model according to forecast date and use only fixed values, we may have improved the score.</li>\n</ul>\n<h1>3. model</h1>\n<ul>\n<li>LSTM (keras)<ul>\n<li>input: features in forecast-day + past-days(to 5 days ago)</li>\n<li>output: multi-output (target1, target2, target3, target4)</li>\n<li>model: Input(6days) &gt; TimeDistributed(Dense) &gt; LSTM &gt; LSTM &gt; Dense(256&gt;128&gt;64&gt;4)</li></ul></li>\n<li>other model:<ul>\n<li>MLP (input: features in only forecast-day)</li>\n<li>GBDT (input: features in only forecast-day)</li></ul></li>\n</ul>\n<p><a href=\"https://postimg.cc/dDwL0JHh\" target=\"_blank\"><img src=\"https://i.postimg.cc/T2WmCwnJ/img1-lstm.png\" alt=\"img1-lstm.png\"></a></p>\n<h1>4. ensemble</h1>\n<ul>\n<li>LSTM:MLP:GBDT = 5:1:1</li>\n<li>score in public LB (evaluation data: 2021-05)<ul>\n<li>LSTM: LB=1.28</li>\n<li>MLP : LB=1.32</li>\n<li>GBDT: LB=1.36</li>\n<li>ensemble: LB=1.26</li></ul></li>\n</ul>\n<h1>5. measures to avoid submission-errors</h1>\n<ul>\n<li>add a lot of exception handling so that it works even if the table or data is missing.</li>\n<li>use same function in both training and predicting</li>\n<li>create dummy data of 2021-08 and 2021-09, and confirmed to work using dummy data on local PC and Kaggle'notebook.</li>\n<li>use API Emulator ( <a href=\"https://www.kaggle.com/nyanpn\" target=\"_blank\">@nyanpn</a> ). Thank you !<br>\n<a href=\"https://www.kaggle.com/nyanpn/api-emulator-for-debugging-your-code-locally\" target=\"_blank\">API Emulator for debugging your code locally</a></li>\n</ul>\n<p>Thank you to my teammate. Let's do it together again!</p>",
  "messages": [
    {
      "id": "1510055",
      "postDate": "09/12/2021 03:40:03",
      "content": "<p>Thank you to the organizers for the fun competition and everyone who participated.<br>\nAnd thank you to my teammate ( taksai <a href=\"https://www.kaggle.com/tsaito21219\" target=\"_blank\">@tsaito21219</a> ).</p>\n<p>I share our team's solution.</p>\n<h1>summary:</h1>\n<ul>\n<li>LSTM model</li>\n<li>valid: 2021-05, 2021-06, 2021-07</li>\n<li>ensemble LSTM:MLP:LightGBM=5:1:1</li>\n</ul>\n<h1>1. validation</h1>\n<ul>\n<li>private-data(2021-08) is in-season, so we use in-season data as validation, and select the month near 2021-08.</li>\n<li>valid: 2021-05, 2021-06, 2021-07</li>\n<li>prepared some patterns as training-data while shifting the period. </li>\n<li>trained 10-fold for each model</li>\n</ul>\n<p><a href=\"https://postimg.cc/nCds5N1C\" target=\"_blank\"><img src=\"https://i.postimg.cc/pLvD9HhY/img1-validation.png\" alt=\"img1-validation.png\"></a></p>\n<h1>2. preprocess data</h1>\n<ul>\n<li>features: total number of 232<ul>\n<li>mean/median/std/min/max of target per playerId last month</li>\n<li>join each table by key(date/playerId/teamId), use the almost feature of each table</li>\n<li>use all tables except events.csv</li>\n<li>not use target-lag-feature</li></ul></li>\n<li>didn't use target-lag-feature, because it's risky. However, if spliting the model according to forecast date and use only fixed values, we may have improved the score.</li>\n</ul>\n<h1>3. model</h1>\n<ul>\n<li>LSTM (keras)<ul>\n<li>input: features in forecast-day + past-days(to 5 days ago)</li>\n<li>output: multi-output (target1, target2, target3, target4)</li>\n<li>model: Input(6days) &gt; TimeDistributed(Dense) &gt; LSTM &gt; LSTM &gt; Dense(256&gt;128&gt;64&gt;4)</li></ul></li>\n<li>other model:<ul>\n<li>MLP (input: features in only forecast-day)</li>\n<li>GBDT (input: features in only forecast-day)</li></ul></li>\n</ul>\n<p><a href=\"https://postimg.cc/dDwL0JHh\" target=\"_blank\"><img src=\"https://i.postimg.cc/T2WmCwnJ/img1-lstm.png\" alt=\"img1-lstm.png\"></a></p>\n<h1>4. ensemble</h1>\n<ul>\n<li>LSTM:MLP:GBDT = 5:1:1</li>\n<li>score in public LB (evaluation data: 2021-05)<ul>\n<li>LSTM: LB=1.28</li>\n<li>MLP : LB=1.32</li>\n<li>GBDT: LB=1.36</li>\n<li>ensemble: LB=1.26</li></ul></li>\n</ul>\n<h1>5. measures to avoid submission-errors</h1>\n<ul>\n<li>add a lot of exception handling so that it works even if the table or data is missing.</li>\n<li>use same function in both training and predicting</li>\n<li>create dummy data of 2021-08 and 2021-09, and confirmed to work using dummy data on local PC and Kaggle'notebook.</li>\n<li>use API Emulator ( <a href=\"https://www.kaggle.com/nyanpn\" target=\"_blank\">@nyanpn</a> ). Thank you !<br>\n<a href=\"https://www.kaggle.com/nyanpn/api-emulator-for-debugging-your-code-locally\" target=\"_blank\">API Emulator for debugging your code locally</a></li>\n</ul>\n<p>Thank you to my teammate. Let's do it together again!</p>",
      "rawMarkdown": "Thank you to the organizers for the fun competition and everyone who participated.\nAnd thank you to my teammate ( taksai @tsaito21219 ).\n\nI share our team's solution.\n\n# summary:\n- LSTM model\n- valid: 2021-05, 2021-06, 2021-07\n- ensemble LSTM:MLP:LightGBM=5:1:1\n\n# 1. validation\n- private-data(2021-08) is in-season, so we use in-season data as validation, and select the month near 2021-08.\n- valid: 2021-05, 2021-06, 2021-07\n- prepared some patterns as training-data while shifting the period. \n- trained 10-fold for each model\n\n[![img1-validation.png](https://i.postimg.cc/pLvD9HhY/img1-validation.png)](https://postimg.cc/nCds5N1C)\n\n# 2. preprocess data\n- features: total number of 232\n    - mean/median/std/min/max of target per playerId last month\n    - join each table by key(date/playerId/teamId), use the almost feature of each table\n    - use all tables except events.csv\n    - not use target-lag-feature\n- didn't use target-lag-feature, because it's risky. However, if spliting the model according to forecast date and use only fixed values, we may have improved the score.\n\n# 3. model\n- LSTM (keras)\n    - input: features in forecast-day + past-days(to 5 days ago)\n    - output: multi-output (target1, target2, target3, target4)\n    - model: Input(6days) > TimeDistributed(Dense) > LSTM > LSTM > Dense(256>128>64>4)\n- other model:\n    - MLP (input: features in only forecast-day)\n    - GBDT (input: features in only forecast-day)\n\n[![img1-lstm.png](https://i.postimg.cc/T2WmCwnJ/img1-lstm.png)](https://postimg.cc/dDwL0JHh)\n\n# 4. ensemble\n- LSTM:MLP:GBDT = 5:1:1\n- score in public LB (evaluation data: 2021-05)\n    - LSTM: LB=1.28\n    - MLP : LB=1.32\n    - GBDT: LB=1.36\n    - ensemble: LB=1.26\n\n# 5. measures to avoid submission-errors\n- add a lot of exception handling so that it works even if the table or data is missing.\n- use same function in both training and predicting\n- create dummy data of 2021-08 and 2021-09, and confirmed to work using dummy data on local PC and Kaggle'notebook.\n- use API Emulator ( @nyanpn ). Thank you !\n[API Emulator for debugging your code locally](https://www.kaggle.com/nyanpn/api-emulator-for-debugging-your-code-locally)\n\nThank you to my teammate. Let's do it together again!",
      "votes": null
    },
    {
      "id": "1511397",
      "postDate": "09/13/2021 11:44:25",
      "content": "<p>Congratulations and well deserved! Using LSTM is a smart idea.</p>",
      "rawMarkdown": "Congratulations and well deserved! Using LSTM is a smart idea.",
      "votes": null
    },
    {
      "id": "1511816",
      "postDate": "09/13/2021 17:54:39",
      "content": "<p>Thank you.  Congratulations to you! <br>\nAPI simulator is very helpful for us. Thanks for sharing.</p>",
      "rawMarkdown": "Thank you.  Congratulations to you! \nAPI simulator is very helpful for us. Thanks for sharing.",
      "votes": null
    },
    {
      "id": "1513269",
      "postDate": "09/15/2021 03:06:28",
      "content": "<p>Congrats! </p>\n<p>Can you share more information regarding how you created the embeddings for the categorical variables?</p>",
      "rawMarkdown": "Congrats! \n\nCan you share more information regarding how you created the embeddings for the categorical variables?",
      "votes": null
    },
    {
      "id": "1513769",
      "postDate": "09/15/2021 11:55:42",
      "content": "<p>I used Embedding of Keras.<br>\n<a href=\"https://keras.io/ja/layers/embeddings/\" target=\"_blank\">link</a></p>\n<ol>\n<li>transform category'name to id (use LabelEncoder)</li>\n<li>convert id to vector using Embeding layer</li>\n<li>combined with numeric layer </li>\n</ol>\n<p>For example. when I have 30 numeric variables and 2 categorical variables, process with code like below.</p>\n<pre><code>input_x = Input(shape=(32, ))\n\n# numeric\nx_num = input_x[:, :30]\nx_num = Dense(20)(x_num)\n\n# category1\nx_cat1 = input_x[:, 30]\nx_embed1 = Embedding(input_dim=100, output_dim=30)(x_cat1)\n\n# category2\nx_cat2 = input_x[:, 31]\nx_embed2 = Embedding(input_dim=200, output_dim=50)(x_cat2)\n\n# concat\nx = Concatenate()([x_num, x_embed1, x_embed2])\n\nout = Dense(4)(x)\n\nmodel = Model(inputs=input_x, outputs=out)\nprint(model.summary())\n</code></pre>",
      "rawMarkdown": "I used Embedding of Keras.\n[link](https://keras.io/ja/layers/embeddings/)\n\n1. transform category'name to id (use LabelEncoder)\n2. convert id to vector using Embeding layer\n3. combined with numeric layer \n\nFor example. when I have 30 numeric variables and 2 categorical variables, process with code like below.\n```\ninput_x = Input(shape=(32, ))\n\n# numeric\nx_num = input_x[:, :30]\nx_num = Dense(20)(x_num)\n\n# category1\nx_cat1 = input_x[:, 30]\nx_embed1 = Embedding(input_dim=100, output_dim=30)(x_cat1)\n\n# category2\nx_cat2 = input_x[:, 31]\nx_embed2 = Embedding(input_dim=200, output_dim=50)(x_cat2)\n\n# concat\nx = Concatenate()([x_num, x_embed1, x_embed2])\n\nout = Dense(4)(x)\n\nmodel = Model(inputs=input_x, outputs=out)\nprint(model.summary())\n```",
      "votes": null
    },
    {
      "id": "1514004",
      "postDate": "09/15/2021 15:38:58",
      "content": "<p>Thank you very much!</p>",
      "rawMarkdown": "Thank you very much!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1511397,
      "author_name": "nyanpn",
      "author_url": "",
      "post_date": "09/13/2021 11:44:25",
      "content": "<p>Congratulations and well deserved! Using LSTM is a smart idea.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1511816,
          "author_name": "moromoromoro",
          "author_url": "",
          "post_date": "09/13/2021 17:54:39",
          "content": "<p>Thank you.  Congratulations to you! <br>\nAPI simulator is very helpful for us. Thanks for sharing.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1513269,
      "author_name": "redfoongus",
      "author_url": "",
      "post_date": "09/15/2021 03:06:28",
      "content": "<p>Congrats! </p>\n<p>Can you share more information regarding how you created the embeddings for the categorical variables?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1513769,
          "author_name": "moromoromoro",
          "author_url": "",
          "post_date": "09/15/2021 11:55:42",
          "content": "<p>I used Embedding of Keras.<br>\n<a href=\"https://keras.io/ja/layers/embeddings/\" target=\"_blank\">link</a></p>\n<ol>\n<li>transform category'name to id (use LabelEncoder)</li>\n<li>convert id to vector using Embeding layer</li>\n<li>combined with numeric layer </li>\n</ol>\n<p>For example. when I have 30 numeric variables and 2 categorical variables, process with code like below.</p>\n<pre><code>input_x = Input(shape=(32, ))\n\n# numeric\nx_num = input_x[:, :30]\nx_num = Dense(20)(x_num)\n\n# category1\nx_cat1 = input_x[:, 30]\nx_embed1 = Embedding(input_dim=100, output_dim=30)(x_cat1)\n\n# category2\nx_cat2 = input_x[:, 31]\nx_embed2 = Embedding(input_dim=200, output_dim=50)(x_cat2)\n\n# concat\nx = Concatenate()([x_num, x_embed1, x_embed2])\n\nout = Dense(4)(x)\n\nmodel = Model(inputs=input_x, outputs=out)\nprint(model.summary())\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1514004,
          "author_name": "redfoongus",
          "author_url": "",
          "post_date": "09/15/2021 15:38:58",
          "content": "<p>Thank you very much!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1510055": "Thank you to the organizers for the fun competition and everyone who participated.\nAnd thank you to my teammate ( taksai @tsaito21219 ).\n\nI share our team's solution.\n\n# summary:\n- LSTM model\n- valid: 2021-05, 2021-06, 2021-07\n- ensemble LSTM:MLP:LightGBM=5:1:1\n\n# 1. validation\n- private-data(2021-08) is in-season, so we use in-season data as validation, and select the month near 2021-08.\n- valid: 2021-05, 2021-06, 2021-07\n- prepared some patterns as training-data while shifting the period. \n- trained 10-fold for each model\n\n[![img1-validation.png](https://i.postimg.cc/pLvD9HhY/img1-validation.png)](https://postimg.cc/nCds5N1C)\n\n# 2. preprocess data\n- features: total number of 232\n    - mean/median/std/min/max of target per playerId last month\n    - join each table by key(date/playerId/teamId), use the almost feature of each table\n    - use all tables except events.csv\n    - not use target-lag-feature\n- didn't use target-lag-feature, because it's risky. However, if spliting the model according to forecast date and use only fixed values, we may have improved the score.\n\n# 3. model\n- LSTM (keras)\n    - input: features in forecast-day + past-days(to 5 days ago)\n    - output: multi-output (target1, target2, target3, target4)\n    - model: Input(6days) > TimeDistributed(Dense) > LSTM > LSTM > Dense(256>128>64>4)\n- other model:\n    - MLP (input: features in only forecast-day)\n    - GBDT (input: features in only forecast-day)\n\n[![img1-lstm.png](https://i.postimg.cc/T2WmCwnJ/img1-lstm.png)](https://postimg.cc/dDwL0JHh)\n\n# 4. ensemble\n- LSTM:MLP:GBDT = 5:1:1\n- score in public LB (evaluation data: 2021-05)\n    - LSTM: LB=1.28\n    - MLP : LB=1.32\n    - GBDT: LB=1.36\n    - ensemble: LB=1.26\n\n# 5. measures to avoid submission-errors\n- add a lot of exception handling so that it works even if the table or data is missing.\n- use same function in both training and predicting\n- create dummy data of 2021-08 and 2021-09, and confirmed to work using dummy data on local PC and Kaggle'notebook.\n- use API Emulator ( @nyanpn ). Thank you !\n[API Emulator for debugging your code locally](https://www.kaggle.com/nyanpn/api-emulator-for-debugging-your-code-locally)\n\nThank you to my teammate. Let's do it together again!",
    "1511397": "Congratulations and well deserved! Using LSTM is a smart idea.",
    "1511816": "Thank you.  Congratulations to you! \nAPI simulator is very helpful for us. Thanks for sharing.",
    "1513269": "Congrats! \n\nCan you share more information regarding how you created the embeddings for the categorical variables?",
    "1513769": "I used Embedding of Keras.\n[link](https://keras.io/ja/layers/embeddings/)\n\n1. transform category'name to id (use LabelEncoder)\n2. convert id to vector using Embeding layer\n3. combined with numeric layer \n\nFor example. when I have 30 numeric variables and 2 categorical variables, process with code like below.\n```\ninput_x = Input(shape=(32, ))\n\n# numeric\nx_num = input_x[:, :30]\nx_num = Dense(20)(x_num)\n\n# category1\nx_cat1 = input_x[:, 30]\nx_embed1 = Embedding(input_dim=100, output_dim=30)(x_cat1)\n\n# category2\nx_cat2 = input_x[:, 31]\nx_embed2 = Embedding(input_dim=200, output_dim=50)(x_cat2)\n\n# concat\nx = Concatenate()([x_num, x_embed1, x_embed2])\n\nout = Dense(4)(x)\n\nmodel = Model(inputs=input_x, outputs=out)\nprint(model.summary())\n```",
    "1514004": "Thank you very much!"
  },
  "source": "meta"
}