{
  "id": 250945,
  "title": "approaches I did not find improvement",
  "url": "/competitions/mlb-player-digital-engagement-forecasting/discussion/250945",
  "author_name": "",
  "post_date": "2021-07-05T07:57:27.637892100Z",
  "votes": 5,
  "comment_count": 14,
  "views": 0,
  "content": "<p>・add playertwitterfollower as a feature.<br>\n・count Nan by row<br>\n・add months as feature<br>\n・target encode some categorical features<br>\n・add Target1~4 of the previous day as a feature<br>\n・add divisionrank and leaguerank as a feature</p>\n<p>It is possible that there were mistakes in my data processing and that these approaches did not work, even though they are inherently effective.<br>\nThese often improved scores on the localCV, but not on the LB.</p>\n<p>I would be happy to hear from you if you report that the approach I just described improved your LB, or if you report an approach that did not work as well as the one I just described.</p>",
  "messages": [
    {
      "id": "1376556",
      "postDate": "07/05/2021 07:57:27",
      "content": "<p>・add playertwitterfollower as a feature.<br>\n・count Nan by row<br>\n・add months as feature<br>\n・target encode some categorical features<br>\n・add Target1~4 of the previous day as a feature<br>\n・add divisionrank and leaguerank as a feature</p>\n<p>It is possible that there were mistakes in my data processing and that these approaches did not work, even though they are inherently effective.<br>\nThese often improved scores on the localCV, but not on the LB.</p>\n<p>I would be happy to hear from you if you report that the approach I just described improved your LB, or if you report an approach that did not work as well as the one I just described.</p>",
      "rawMarkdown": "・add playertwitterfollower as a feature.\n・count Nan by row\n・add months as feature\n・target encode some categorical features\n・add Target1~4 of the previous day as a feature\n・add divisionrank and leaguerank as a feature\n\nIt is possible that there were mistakes in my data processing and that these approaches did not work, even though they are inherently effective.\nThese often improved scores on the localCV, but not on the LB.\n\nI would be happy to hear from you if you report that the approach I just described improved your LB, or if you report an approach that did not work as well as the one I just described.",
      "votes": null
    },
    {
      "id": "1376559",
      "postDate": "07/05/2021 08:01:09",
      "content": "<p>P.S.. The reason I posted this discussion is that I thought that features that I judged to be ineffective might be effective with proper processing.</p>",
      "rawMarkdown": "P.S.. The reason I posted this discussion is that I thought that features that I judged to be ineffective might be effective with proper processing.",
      "votes": null
    },
    {
      "id": "1376631",
      "postDate": "07/05/2021 08:33:53",
      "content": "<ul>\n<li>month feature didnt improve for me also.</li>\n<li>target features neither</li>\n<li>playerTwitter followers did improve a bit my cv</li>\n</ul>",
      "rawMarkdown": "month feature didnt improve for me also.\n- target features neither\n- playerTwitter followers did improve a bit my cv",
      "votes": null
    },
    {
      "id": "1379266",
      "postDate": "07/07/2021 08:03:50",
      "content": "<p>Twitter followers and awards data did not help me either, although Twitter followers helped CV a bit.  Player salaries as well did not help.</p>",
      "rawMarkdown": "Twitter followers and awards data did not help me either, although Twitter followers helped CV a bit.  Player salaries as well did not help.",
      "votes": null
    },
    {
      "id": "1379782",
      "postDate": "07/07/2021 15:37:04",
      "content": "<p>Same with me.</p>",
      "rawMarkdown": "Same with me.",
      "votes": null
    },
    {
      "id": "1383850",
      "postDate": "07/11/2021 09:26:07",
      "content": "<p>\"add Target1~4 of the previous day as a feature\" works very nicely if we use more old histories.<br>\nvalidation score enhanced:<br>\nprovide -1 day -&gt; -3% <br>\nprovide -2 day -&gt; -1% <br>\nprovide -3 day -&gt; -0.001% <br>\nprovide -4 day -&gt; +0.7% <br>\nprovide -5 day -&gt; +0.9% <br>\nprovide -6 day -&gt; +0.4% <br>\nprovide -7 day -&gt; +0.2% <br>\nprovide -10 day -&gt; +0.1% <br>\nprovide -30 day -&gt; -0.0% <br>\nprovide -365 day -&gt; -0.033%<br>\n\"add playertwitterfollower as a feature.\" also work with percentile normalization,<br>\nimprove 1.2%, LB 1.3433 -&gt; LB 1.3336</p>\n<p>can I ask your testing parameters?<br>\nsuch as, lgbm boosting method, max depth, leaves, and processor count?</p>",
      "rawMarkdown": "\"add Target1~4 of the previous day as a feature\" works very nicely if we use more old histories.\nvalidation score enhanced:\nprovide -1 day -> -3% \nprovide -2 day -> -1% \nprovide -3 day -> -0.001% \nprovide -4 day -> +0.7% \nprovide -5 day -> +0.9% \nprovide -6 day -> +0.4% \nprovide -7 day -> +0.2% \nprovide -10 day -> +0.1% \nprovide -30 day -> -0.0% \nprovide -365 day -> -0.033%\n\"add playertwitterfollower as a feature.\" also work with percentile normalization,\nimprove 1.2%, LB 1.3433 -> LB 1.3336\n\ncan I ask your testing parameters?\nsuch as, lgbm boosting method, max depth, leaves, and processor count?",
      "votes": null
    },
    {
      "id": "1383891",
      "postDate": "07/11/2021 10:08:10",
      "content": "<p>I am testing lgbm model based on this kernel.<br>\n<a href=\"https://www.kaggle.com/columbia2131/mlb-lightgbm-starter-dataset-code-en-ja\" target=\"_blank\">https://www.kaggle.com/columbia2131/mlb-lightgbm-starter-dataset-code-en-ja</a>  <br>\nI use  same params as in this kernel.　<br>\n<a href=\"https://www.kaggle.com/mlconsult/1-35-lightgbm-ann\" target=\"_blank\">https://www.kaggle.com/mlconsult/1-35-lightgbm-ann</a><br>\nI guess that the difference in the impact of adding Twitter followers to the features I believe that the difference in the impact of adding Twitter followers to the model is caused by the difference in missing value processing.</p>",
      "rawMarkdown": "I am testing lgbm model based on this kernel.\nhttps://www.kaggle.com/columbia2131/mlb-lightgbm-starter-dataset-code-en-ja  \nI use  same params as in this kernel.　\nhttps://www.kaggle.com/mlconsult/1-35-lightgbm-ann\nI guess that the difference in the impact of adding Twitter followers to the features I believe that the difference in the impact of adding Twitter followers to the model is caused by the difference in missing value processing.",
      "votes": null
    },
    {
      "id": "1383892",
      "postDate": "07/11/2021 10:08:18",
      "content": "<p>How do you provide those lags on private test? With the las predictions/submissions? </p>",
      "rawMarkdown": "How do you provide those lags on private test? With the las predictions/submissions?",
      "votes": null
    },
    {
      "id": "1383907",
      "postDate": "07/11/2021 10:18:20",
      "content": "<p><a href=\"https://www.kaggle.com/enric1296\" target=\"_blank\">@enric1296</a> lags are generate by model's prediction if input was provided by test .<br>\nif it was provided in train.csv, then use it. otherwise use null or predicted (or submitted) result.</p>",
      "rawMarkdown": "enric1296 lags are generate by model's prediction if input was provided by test .\nif it was provided in train.csv, then use it. otherwise use null or predicted (or submitted) result.",
      "votes": null
    },
    {
      "id": "1383910",
      "postDate": "07/11/2021 10:24:02",
      "content": "<p>Past targets are provided by test? I guess they are not and you need to use past predictions as feature lags</p>",
      "rawMarkdown": "Past targets are provided by test? I guess they are not and you need to use past predictions as feature lags",
      "votes": null
    },
    {
      "id": "1383918",
      "postDate": "07/11/2021 10:35:53",
      "content": "<p><a href=\"https://www.kaggle.com/enric1296\" target=\"_blank\">@enric1296</a> I think I said vaguely. There is no target. What I was talking is the case of providing information (except targets) for inference.</p>",
      "rawMarkdown": "enric1296 I think I said vaguely. There is no target. What I was talking is the case of providing information (except targets) for inference.",
      "votes": null
    },
    {
      "id": "1383939",
      "postDate": "07/11/2021 11:05:42",
      "content": "<p>Oh now ive understood, train features for test as the  test dates are a continuation of the train ones and in order to generate lag features you save past predictions in addition to train ones. Is it correct? </p>",
      "rawMarkdown": "Oh now ive understood, train features for test as the  test dates are a continuation of the train ones and in order to generate lag features you save past predictions in addition to train ones. Is it correct?",
      "votes": null
    },
    {
      "id": "1384422",
      "postDate": "07/11/2021 19:57:29",
      "content": "<p>Supposedly the test set is not necessarily continuous, doesn't this run into the risk of having a lot of nan's in that scenario?</p>",
      "rawMarkdown": "Supposedly the test set is not necessarily continuous, doesn't this run into the risk of having a lot of nan's in that scenario?",
      "votes": null
    },
    {
      "id": "1385745",
      "postDate": "07/13/2021 02:30:08",
      "content": "<p>My good friend, my email is ghj_xin@163.com, how can I contact you?</p>",
      "rawMarkdown": "My good friend, my email is ghj_xin@163.com, how can I contact you?",
      "votes": null
    },
    {
      "id": "1385748",
      "postDate": "07/13/2021 02:37:23",
      "content": "<p>I know!I'm very pleased!</p>",
      "rawMarkdown": "I know!I'm very pleased!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1376559,
      "author_name": "sinrinosekai",
      "author_url": "",
      "post_date": "07/05/2021 08:01:09",
      "content": "<p>P.S.. The reason I posted this discussion is that I thought that features that I judged to be ineffective might be effective with proper processing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1376631,
      "author_name": "enric1296",
      "author_url": "",
      "post_date": "07/05/2021 08:33:53",
      "content": "<ul>\n<li>month feature didnt improve for me also.</li>\n<li>target features neither</li>\n<li>playerTwitter followers did improve a bit my cv</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1379266,
      "author_name": "chinta",
      "author_url": "",
      "post_date": "07/07/2021 08:03:50",
      "content": "<p>Twitter followers and awards data did not help me either, although Twitter followers helped CV a bit.  Player salaries as well did not help.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1379782,
          "author_name": "joshspchang",
          "author_url": "",
          "post_date": "07/07/2021 15:37:04",
          "content": "<p>Same with me.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1383850,
      "author_name": "assign",
      "author_url": "",
      "post_date": "07/11/2021 09:26:07",
      "content": "<p>\"add Target1~4 of the previous day as a feature\" works very nicely if we use more old histories.<br>\nvalidation score enhanced:<br>\nprovide -1 day -&gt; -3% <br>\nprovide -2 day -&gt; -1% <br>\nprovide -3 day -&gt; -0.001% <br>\nprovide -4 day -&gt; +0.7% <br>\nprovide -5 day -&gt; +0.9% <br>\nprovide -6 day -&gt; +0.4% <br>\nprovide -7 day -&gt; +0.2% <br>\nprovide -10 day -&gt; +0.1% <br>\nprovide -30 day -&gt; -0.0% <br>\nprovide -365 day -&gt; -0.033%<br>\n\"add playertwitterfollower as a feature.\" also work with percentile normalization,<br>\nimprove 1.2%, LB 1.3433 -&gt; LB 1.3336</p>\n<p>can I ask your testing parameters?<br>\nsuch as, lgbm boosting method, max depth, leaves, and processor count?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1383891,
          "author_name": "sinrinosekai",
          "author_url": "",
          "post_date": "07/11/2021 10:08:10",
          "content": "<p>I am testing lgbm model based on this kernel.<br>\n<a href=\"https://www.kaggle.com/columbia2131/mlb-lightgbm-starter-dataset-code-en-ja\" target=\"_blank\">https://www.kaggle.com/columbia2131/mlb-lightgbm-starter-dataset-code-en-ja</a>  <br>\nI use  same params as in this kernel.　<br>\n<a href=\"https://www.kaggle.com/mlconsult/1-35-lightgbm-ann\" target=\"_blank\">https://www.kaggle.com/mlconsult/1-35-lightgbm-ann</a><br>\nI guess that the difference in the impact of adding Twitter followers to the features I believe that the difference in the impact of adding Twitter followers to the model is caused by the difference in missing value processing.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1383892,
          "author_name": "enric1296",
          "author_url": "",
          "post_date": "07/11/2021 10:08:18",
          "content": "<p>How do you provide those lags on private test? With the las predictions/submissions? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1383907,
          "author_name": "assign",
          "author_url": "",
          "post_date": "07/11/2021 10:18:20",
          "content": "<p><a href=\"https://www.kaggle.com/enric1296\" target=\"_blank\">@enric1296</a> lags are generate by model's prediction if input was provided by test .<br>\nif it was provided in train.csv, then use it. otherwise use null or predicted (or submitted) result.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1383910,
          "author_name": "enric1296",
          "author_url": "",
          "post_date": "07/11/2021 10:24:02",
          "content": "<p>Past targets are provided by test? I guess they are not and you need to use past predictions as feature lags</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1383918,
          "author_name": "assign",
          "author_url": "",
          "post_date": "07/11/2021 10:35:53",
          "content": "<p><a href=\"https://www.kaggle.com/enric1296\" target=\"_blank\">@enric1296</a> I think I said vaguely. There is no target. What I was talking is the case of providing information (except targets) for inference.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1383939,
          "author_name": "enric1296",
          "author_url": "",
          "post_date": "07/11/2021 11:05:42",
          "content": "<p>Oh now ive understood, train features for test as the  test dates are a continuation of the train ones and in order to generate lag features you save past predictions in addition to train ones. Is it correct? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1384422,
          "author_name": "redfoongus",
          "author_url": "",
          "post_date": "07/11/2021 19:57:29",
          "content": "<p>Supposedly the test set is not necessarily continuous, doesn't this run into the risk of having a lot of nan's in that scenario?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1385745,
      "author_name": "guohey",
      "author_url": "",
      "post_date": "07/13/2021 02:30:08",
      "content": "<p>My good friend, my email is ghj_xin@163.com, how can I contact you?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1385748,
      "author_name": "guohey",
      "author_url": "",
      "post_date": "07/13/2021 02:37:23",
      "content": "<p>I know!I'm very pleased!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1376556": "・add playertwitterfollower as a feature.\n・count Nan by row\n・add months as feature\n・target encode some categorical features\n・add Target1~4 of the previous day as a feature\n・add divisionrank and leaguerank as a feature\n\nIt is possible that there were mistakes in my data processing and that these approaches did not work, even though they are inherently effective.\nThese often improved scores on the localCV, but not on the LB.\n\nI would be happy to hear from you if you report that the approach I just described improved your LB, or if you report an approach that did not work as well as the one I just described.",
    "1376559": "P.S.. The reason I posted this discussion is that I thought that features that I judged to be ineffective might be effective with proper processing.",
    "1376631": "month feature didnt improve for me also.\n- target features neither\n- playerTwitter followers did improve a bit my cv",
    "1379266": "Twitter followers and awards data did not help me either, although Twitter followers helped CV a bit.  Player salaries as well did not help.",
    "1379782": "Same with me.",
    "1383850": "\"add Target1~4 of the previous day as a feature\" works very nicely if we use more old histories.\nvalidation score enhanced:\nprovide -1 day -> -3% \nprovide -2 day -> -1% \nprovide -3 day -> -0.001% \nprovide -4 day -> +0.7% \nprovide -5 day -> +0.9% \nprovide -6 day -> +0.4% \nprovide -7 day -> +0.2% \nprovide -10 day -> +0.1% \nprovide -30 day -> -0.0% \nprovide -365 day -> -0.033%\n\"add playertwitterfollower as a feature.\" also work with percentile normalization,\nimprove 1.2%, LB 1.3433 -> LB 1.3336\n\ncan I ask your testing parameters?\nsuch as, lgbm boosting method, max depth, leaves, and processor count?",
    "1383891": "I am testing lgbm model based on this kernel.\nhttps://www.kaggle.com/columbia2131/mlb-lightgbm-starter-dataset-code-en-ja  \nI use  same params as in this kernel.　\nhttps://www.kaggle.com/mlconsult/1-35-lightgbm-ann\nI guess that the difference in the impact of adding Twitter followers to the features I believe that the difference in the impact of adding Twitter followers to the model is caused by the difference in missing value processing.",
    "1383892": "How do you provide those lags on private test? With the las predictions/submissions?",
    "1383907": "enric1296 lags are generate by model's prediction if input was provided by test .\nif it was provided in train.csv, then use it. otherwise use null or predicted (or submitted) result.",
    "1383910": "Past targets are provided by test? I guess they are not and you need to use past predictions as feature lags",
    "1383918": "enric1296 I think I said vaguely. There is no target. What I was talking is the case of providing information (except targets) for inference.",
    "1383939": "Oh now ive understood, train features for test as the  test dates are a continuation of the train ones and in order to generate lag features you save past predictions in addition to train ones. Is it correct?",
    "1384422": "Supposedly the test set is not necessarily continuous, doesn't this run into the risk of having a lot of nan's in that scenario?",
    "1385745": "My good friend, my email is ghj_xin@163.com, how can I contact you?",
    "1385748": "I know!I'm very pleased!"
  },
  "source": "meta"
}