{
  "id": 253940,
  "title": "About the targets",
  "url": "/competitions/mlb-player-digital-engagement-forecasting/discussion/253940",
  "author_name": "",
  "post_date": "2021-07-19T13:25:30.101326700Z",
  "votes": 41,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I think we all realized that the targets are capped at 100. This is probably due to max scaling. And the max scaling is applied to each player individually. Assuming this and assuming that the minimum non-zero target value for the players with very low target values should be one unit, one can figure out the max scaling constant. </p>\n<p><img src=\"https://i.imgur.com/GFsmyaB.png\" alt=\"\"></p>\n<p>target1 for player 425772 is divided by 663077, while target1 for player 425784 is divided by 1753059. Getting such exact numbers proves the theory. This also explains why features like total Twitter followers are not useful since each player is normalized by its max.</p>\n<p>Given that the competition host is Google and Google Trends are always max scaled individually, the targets are probably the search trends. There are 5 different trends categories:</p>\n<ul>\n<li>Web Search</li>\n<li>Image Search</li>\n<li>News Search</li>\n<li>Google Shopping Search</li>\n<li>YouTube Search</li>\n</ul>\n<p>Since we have only 4 targets, one of these targets might be not used due to its low volume, which may be the shopping one.</p>\n<p>What happens to the players with the same name then? I don't know how they do it but Google Trends is able to distinguish these players. If you look at Will Smith example, Google Trends have 3 options: actor, catcher and pitcher. <a href=\"https://trends.google.com/trends/explore?date=2018-01-01%202021-07-19&amp;geo=US&amp;q=%2Fm%2F0jw99sx\" target=\"_blank\">https://trends.google.com/trends/explore?date=2018-01-01%202021-07-19&amp;geo=US&amp;q=%2Fm%2F0jw99sx</a></p>\n<p>I think they preferred to not reveal the targets to protect the public LB, so that people don't start using May 2021 Google Trends data to game the system. But maybe some people are already doing it.</p>\n<p>The most important question is how Kaggle does the max normalization for the upcoming test set. Options:</p>\n<ul>\n<li>The divisor stays the same, if the upcoming data have larger values than the divisor we may have target values larger than 100.</li>\n<li>The divisor stays the same but values larger than 100 are clipped to 100.</li>\n<li>The divisor is updated every day with the new data. The new maximum becomes the new divisor. Then I wonder if the new divisor will affect the train set.</li>\n</ul>\n<p>I hope Kaggle makes it clear about which normalization technique will be used. In real life problems, we know how the data is normalized and knowing it here will make the competition more realistic. Otherwise, people may gamble with several assumptions.</p>",
  "messages": [
    {
      "id": "1393223",
      "postDate": "07/19/2021 13:25:30",
      "content": "<p>I think we all realized that the targets are capped at 100. This is probably due to max scaling. And the max scaling is applied to each player individually. Assuming this and assuming that the minimum non-zero target value for the players with very low target values should be one unit, one can figure out the max scaling constant. </p>\n<p><img src=\"https://i.imgur.com/GFsmyaB.png\" alt=\"\"></p>\n<p>target1 for player 425772 is divided by 663077, while target1 for player 425784 is divided by 1753059. Getting such exact numbers proves the theory. This also explains why features like total Twitter followers are not useful since each player is normalized by its max.</p>\n<p>Given that the competition host is Google and Google Trends are always max scaled individually, the targets are probably the search trends. There are 5 different trends categories:</p>\n<ul>\n<li>Web Search</li>\n<li>Image Search</li>\n<li>News Search</li>\n<li>Google Shopping Search</li>\n<li>YouTube Search</li>\n</ul>\n<p>Since we have only 4 targets, one of these targets might be not used due to its low volume, which may be the shopping one.</p>\n<p>What happens to the players with the same name then? I don't know how they do it but Google Trends is able to distinguish these players. If you look at Will Smith example, Google Trends have 3 options: actor, catcher and pitcher. <a href=\"https://trends.google.com/trends/explore?date=2018-01-01%202021-07-19&amp;geo=US&amp;q=%2Fm%2F0jw99sx\" target=\"_blank\">https://trends.google.com/trends/explore?date=2018-01-01%202021-07-19&amp;geo=US&amp;q=%2Fm%2F0jw99sx</a></p>\n<p>I think they preferred to not reveal the targets to protect the public LB, so that people don't start using May 2021 Google Trends data to game the system. But maybe some people are already doing it.</p>\n<p>The most important question is how Kaggle does the max normalization for the upcoming test set. Options:</p>\n<ul>\n<li>The divisor stays the same, if the upcoming data have larger values than the divisor we may have target values larger than 100.</li>\n<li>The divisor stays the same but values larger than 100 are clipped to 100.</li>\n<li>The divisor is updated every day with the new data. The new maximum becomes the new divisor. Then I wonder if the new divisor will affect the train set.</li>\n</ul>\n<p>I hope Kaggle makes it clear about which normalization technique will be used. In real life problems, we know how the data is normalized and knowing it here will make the competition more realistic. Otherwise, people may gamble with several assumptions.</p>",
      "rawMarkdown": "I think we all realized that the targets are capped at 100. This is probably due to max scaling. And the max scaling is applied to each player individually. Assuming this and assuming that the minimum non-zero target value for the players with very low target values should be one unit, one can figure out the max scaling constant. \n\n![](https://i.imgur.com/GFsmyaB.png)\n\ntarget1 for player 425772 is divided by 663077, while target1 for player 425784 is divided by 1753059. Getting such exact numbers proves the theory. This also explains why features like total Twitter followers are not useful since each player is normalized by its max.\n\nGiven that the competition host is Google and Google Trends are always max scaled individually, the targets are probably the search trends. There are 5 different trends categories:\n* Web Search\n* Image Search\n* News Search\n* Google Shopping Search\n* YouTube Search\n\nSince we have only 4 targets, one of these targets might be not used due to its low volume, which may be the shopping one.\n\nWhat happens to the players with the same name then? I don't know how they do it but Google Trends is able to distinguish these players. If you look at Will Smith example, Google Trends have 3 options: actor, catcher and pitcher. https://trends.google.com/trends/explore?date=2018-01-01%202021-07-19&geo=US&q=%2Fm%2F0jw99sx\n\nI think they preferred to not reveal the targets to protect the public LB, so that people don't start using May 2021 Google Trends data to game the system. But maybe some people are already doing it.\n\nThe most important question is how Kaggle does the max normalization for the upcoming test set. Options:\n* The divisor stays the same, if the upcoming data have larger values than the divisor we may have target values larger than 100.\n* The divisor stays the same but values larger than 100 are clipped to 100.\n* The divisor is updated every day with the new data. The new maximum becomes the new divisor. Then I wonder if the new divisor will affect the train set.\n\nI hope Kaggle makes it clear about which normalization technique will be used. In real life problems, we know how the data is normalized and knowing it here will make the competition more realistic. Otherwise, people may gamble with several assumptions.",
      "votes": null
    },
    {
      "id": "1393925",
      "postDate": "07/20/2021 03:05:43",
      "content": "<p>I think that max scaling is applied on a daily basis. In the train data, there is only one data per day which target is 100.</p>",
      "rawMarkdown": "I think that max scaling is applied on a daily basis. In the train data, there is only one data per day which target is 100.",
      "votes": null
    },
    {
      "id": "1394093",
      "postDate": "07/20/2021 06:56:00",
      "content": "<p>This is correct. The only exception is that there are 2 days for <code>target3</code> where there are 2 players with 100. There are also many players that have many days with target==100 which would be unlikely to happen if it was a simple Min-Max scale by player since they'd have to have the exact same unscaled targets on multiple days. </p>",
      "rawMarkdown": "This is correct. The only exception is that there are 2 days for `target3` where there are 2 players with 100. There are also many players that have many days with target==100 which would be unlikely to happen if it was a simple Min-Max scale by player since they'd have to have the exact same unscaled targets on multiple days.",
      "votes": null
    },
    {
      "id": "1394138",
      "postDate": "07/20/2021 07:43:52",
      "content": "<p>Makes sense. This also explains why I don't see the same exact pattern as in Google Trends.</p>",
      "rawMarkdown": "Makes sense. This also explains why I don't see the same exact pattern as in Google Trends.",
      "votes": null
    },
    {
      "id": "1394645",
      "postDate": "07/20/2021 13:38:56",
      "content": "<p>With respect to Twitter, there are 2 sets of followers  - player and team. If these targets have any relationship to them for metrics e.g. likes, retweets, etc. The targets with 100 could be the max value for that day and the rest proportional to it.  So where many players have 100 over multiple days that could be team related as well.  <br>\nHave not tried too much to figure out what the targets represent but what in the data supports high targets. Having found so many issues with dates though not even sure that is possible without a crystal ball. </p>",
      "rawMarkdown": "With respect to Twitter, there are 2 sets of followers  - player and team. If these targets have any relationship to them for metrics e.g. likes, retweets, etc. The targets with 100 could be the max value for that day and the rest proportional to it.  So where many players have 100 over multiple days that could be team related as well.  \nHave not tried too much to figure out what the targets represent but what in the data supports high targets. Having found so many issues with dates though not even sure that is possible without a crystal ball.",
      "votes": null
    },
    {
      "id": "1395431",
      "postDate": "07/21/2021 08:01:21",
      "content": "<p>The data is extremely noisy then. One player noises the whole scale each day. A player could have a good day at his game but get low media engagement just because another player shows up on magazines with some scandal. I don't know if this noise is really intended and why Kaggle didn't run the competition on the exact numbers. Predicting x could be more useful than predicting x/max(x). <a href=\"https://www.kaggle.com/wcukierski\" target=\"_blank\">@wcukierski</a> <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a></p>",
      "rawMarkdown": "The data is extremely noisy then. One player noises the whole scale each day. A player could have a good day at his game but get low media engagement just because another player shows up on magazines with some scandal. I don't know if this noise is really intended and why Kaggle didn't run the competition on the exact numbers. Predicting x could be more useful than predicting x/max(x). @wcukierski @juliaelliott",
      "votes": null
    },
    {
      "id": "1395504",
      "postDate": "07/21/2021 09:14:15",
      "content": "<p>I hope they change the how this competition works. From my local validation, the predictions are absolute garbage (negative r2 score). Aside from how they do scaling, this is probably also because many factors other than in-games stats are at play and that kind of information is simply missing</p>",
      "rawMarkdown": "I hope they change the how this competition works. From my local validation, the predictions are absolute garbage (negative r2 score). Aside from how they do scaling, this is probably also because many factors other than in-games stats are at play and that kind of information is simply missing",
      "votes": null
    },
    {
      "id": "1395568",
      "postDate": "07/21/2021 10:19:41",
      "content": "<p>Multiple 100s for a target could be, data normalization by Geography, there could be two different popular players on a day in two different Geographies.</p>",
      "rawMarkdown": "Multiple 100s for a target could be, data normalization by Geography, there could be two different popular players on a day in two different Geographies.",
      "votes": null
    },
    {
      "id": "1396408",
      "postDate": "07/22/2021 05:04:33",
      "content": "<p><a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> - indeed. The example raised early on of  <a href=\"https://www.kaggle.com/c/mlb-player-digital-engagement-forecasting/discussion/246499\" target=\"_blank\">Tyler Skaggs</a> who died in July 2019 still having targets increasing even in 2021. That had playerForTestSetAndFuturePreds  False so should not be included.<br>\nBut what it showed was if players have other accounts like for Foundations, fund raising activities or are in the news for something like in that instance when staffers were charged and another on painkillers and his death, there were spikes in targets. None of that would be in the competition data. </p>\n<p>My thinking is that the analytics side of this for the explainability prize is the only way to focus on those outside factors.  If there had been more clarity on use of external data and that data could be updated throughout the evaluation phase it might have helped.</p>\n<p>Where the competition data fails also is in transaction dates that are after the fact or rosters are wrong, so players not on rosters are playing, trades, retires, call ups are announced and the spike shows the day previous. OK could be the rumour mill at work but just seems like the dates are wrong.</p>\n<p>Am waiting to see how the updated train to 20 Jul looks. </p>",
      "rawMarkdown": "shujun717 - indeed. The example raised early on of  [Tyler Skaggs](https://www.kaggle.com/c/mlb-player-digital-engagement-forecasting/discussion/246499) who died in July 2019 still having targets increasing even in 2021. That had playerForTestSetAndFuturePreds  False so should not be included.\nBut what it showed was if players have other accounts like for Foundations, fund raising activities or are in the news for something like in that instance when staffers were charged and another on painkillers and his death, there were spikes in targets. None of that would be in the competition data. \n\nMy thinking is that the analytics side of this for the explainability prize is the only way to focus on those outside factors.  If there had been more clarity on use of external data and that data could be updated throughout the evaluation phase it might have helped.\n\nWhere the competition data fails also is in transaction dates that are after the fact or rosters are wrong, so players not on rosters are playing, trades, retires, call ups are announced and the spike shows the day previous. OK could be the rumour mill at work but just seems like the dates are wrong.\n\nAm waiting to see how the updated train to 20 Jul looks.",
      "votes": null
    },
    {
      "id": "1397048",
      "postDate": "07/22/2021 17:57:12",
      "content": "<p>FYI - the promised \"opt-in\" 20 Jul data has been uploaded as <code>train_updated.csv</code> on the <a href=\"https://www.kaggle.com/c/mlb-player-digital-engagement-forecasting/data\" target=\"_blank\">data page</a></p>",
      "rawMarkdown": "FYI - the promised \"opt-in\" 20 Jul data has been uploaded as `train_updated.csv` on the [data page](https://www.kaggle.com/c/mlb-player-digital-engagement-forecasting/data)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1393925,
      "author_name": "masatomatsui",
      "author_url": "",
      "post_date": "07/20/2021 03:05:43",
      "content": "<p>I think that max scaling is applied on a daily basis. In the train data, there is only one data per day which target is 100.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1394093,
          "author_name": "brandenkmurray",
          "author_url": "",
          "post_date": "07/20/2021 06:56:00",
          "content": "<p>This is correct. The only exception is that there are 2 days for <code>target3</code> where there are 2 players with 100. There are also many players that have many days with target==100 which would be unlikely to happen if it was a simple Min-Max scale by player since they'd have to have the exact same unscaled targets on multiple days. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1394138,
          "author_name": "aerdem4",
          "author_url": "",
          "post_date": "07/20/2021 07:43:52",
          "content": "<p>Makes sense. This also explains why I don't see the same exact pattern as in Google Trends.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1394645,
          "author_name": "something4kag",
          "author_url": "",
          "post_date": "07/20/2021 13:38:56",
          "content": "<p>With respect to Twitter, there are 2 sets of followers  - player and team. If these targets have any relationship to them for metrics e.g. likes, retweets, etc. The targets with 100 could be the max value for that day and the rest proportional to it.  So where many players have 100 over multiple days that could be team related as well.  <br>\nHave not tried too much to figure out what the targets represent but what in the data supports high targets. Having found so many issues with dates though not even sure that is possible without a crystal ball. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1395431,
          "author_name": "aerdem4",
          "author_url": "",
          "post_date": "07/21/2021 08:01:21",
          "content": "<p>The data is extremely noisy then. One player noises the whole scale each day. A player could have a good day at his game but get low media engagement just because another player shows up on magazines with some scandal. I don't know if this noise is really intended and why Kaggle didn't run the competition on the exact numbers. Predicting x could be more useful than predicting x/max(x). <a href=\"https://www.kaggle.com/wcukierski\" target=\"_blank\">@wcukierski</a> <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1395504,
          "author_name": "shujun717",
          "author_url": "",
          "post_date": "07/21/2021 09:14:15",
          "content": "<p>I hope they change the how this competition works. From my local validation, the predictions are absolute garbage (negative r2 score). Aside from how they do scaling, this is probably also because many factors other than in-games stats are at play and that kind of information is simply missing</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1395568,
          "author_name": "chinta",
          "author_url": "",
          "post_date": "07/21/2021 10:19:41",
          "content": "<p>Multiple 100s for a target could be, data normalization by Geography, there could be two different popular players on a day in two different Geographies.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1396408,
          "author_name": "something4kag",
          "author_url": "",
          "post_date": "07/22/2021 05:04:33",
          "content": "<p><a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> - indeed. The example raised early on of  <a href=\"https://www.kaggle.com/c/mlb-player-digital-engagement-forecasting/discussion/246499\" target=\"_blank\">Tyler Skaggs</a> who died in July 2019 still having targets increasing even in 2021. That had playerForTestSetAndFuturePreds  False so should not be included.<br>\nBut what it showed was if players have other accounts like for Foundations, fund raising activities or are in the news for something like in that instance when staffers were charged and another on painkillers and his death, there were spikes in targets. None of that would be in the competition data. </p>\n<p>My thinking is that the analytics side of this for the explainability prize is the only way to focus on those outside factors.  If there had been more clarity on use of external data and that data could be updated throughout the evaluation phase it might have helped.</p>\n<p>Where the competition data fails also is in transaction dates that are after the fact or rosters are wrong, so players not on rosters are playing, trades, retires, call ups are announced and the spike shows the day previous. OK could be the rumour mill at work but just seems like the dates are wrong.</p>\n<p>Am waiting to see how the updated train to 20 Jul looks. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1397048,
          "author_name": "juliaelliott",
          "author_url": "",
          "post_date": "07/22/2021 17:57:12",
          "content": "<p>FYI - the promised \"opt-in\" 20 Jul data has been uploaded as <code>train_updated.csv</code> on the <a href=\"https://www.kaggle.com/c/mlb-player-digital-engagement-forecasting/data\" target=\"_blank\">data page</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1393223": "I think we all realized that the targets are capped at 100. This is probably due to max scaling. And the max scaling is applied to each player individually. Assuming this and assuming that the minimum non-zero target value for the players with very low target values should be one unit, one can figure out the max scaling constant. \n\n![](https://i.imgur.com/GFsmyaB.png)\n\ntarget1 for player 425772 is divided by 663077, while target1 for player 425784 is divided by 1753059. Getting such exact numbers proves the theory. This also explains why features like total Twitter followers are not useful since each player is normalized by its max.\n\nGiven that the competition host is Google and Google Trends are always max scaled individually, the targets are probably the search trends. There are 5 different trends categories:\n* Web Search\n* Image Search\n* News Search\n* Google Shopping Search\n* YouTube Search\n\nSince we have only 4 targets, one of these targets might be not used due to its low volume, which may be the shopping one.\n\nWhat happens to the players with the same name then? I don't know how they do it but Google Trends is able to distinguish these players. If you look at Will Smith example, Google Trends have 3 options: actor, catcher and pitcher. https://trends.google.com/trends/explore?date=2018-01-01%202021-07-19&geo=US&q=%2Fm%2F0jw99sx\n\nI think they preferred to not reveal the targets to protect the public LB, so that people don't start using May 2021 Google Trends data to game the system. But maybe some people are already doing it.\n\nThe most important question is how Kaggle does the max normalization for the upcoming test set. Options:\n* The divisor stays the same, if the upcoming data have larger values than the divisor we may have target values larger than 100.\n* The divisor stays the same but values larger than 100 are clipped to 100.\n* The divisor is updated every day with the new data. The new maximum becomes the new divisor. Then I wonder if the new divisor will affect the train set.\n\nI hope Kaggle makes it clear about which normalization technique will be used. In real life problems, we know how the data is normalized and knowing it here will make the competition more realistic. Otherwise, people may gamble with several assumptions.",
    "1393925": "I think that max scaling is applied on a daily basis. In the train data, there is only one data per day which target is 100.",
    "1394093": "This is correct. The only exception is that there are 2 days for `target3` where there are 2 players with 100. There are also many players that have many days with target==100 which would be unlikely to happen if it was a simple Min-Max scale by player since they'd have to have the exact same unscaled targets on multiple days.",
    "1394138": "Makes sense. This also explains why I don't see the same exact pattern as in Google Trends.",
    "1394645": "With respect to Twitter, there are 2 sets of followers  - player and team. If these targets have any relationship to them for metrics e.g. likes, retweets, etc. The targets with 100 could be the max value for that day and the rest proportional to it.  So where many players have 100 over multiple days that could be team related as well.  \nHave not tried too much to figure out what the targets represent but what in the data supports high targets. Having found so many issues with dates though not even sure that is possible without a crystal ball.",
    "1395431": "The data is extremely noisy then. One player noises the whole scale each day. A player could have a good day at his game but get low media engagement just because another player shows up on magazines with some scandal. I don't know if this noise is really intended and why Kaggle didn't run the competition on the exact numbers. Predicting x could be more useful than predicting x/max(x). @wcukierski @juliaelliott",
    "1395504": "I hope they change the how this competition works. From my local validation, the predictions are absolute garbage (negative r2 score). Aside from how they do scaling, this is probably also because many factors other than in-games stats are at play and that kind of information is simply missing",
    "1395568": "Multiple 100s for a target could be, data normalization by Geography, there could be two different popular players on a day in two different Geographies.",
    "1396408": "shujun717 - indeed. The example raised early on of  [Tyler Skaggs](https://www.kaggle.com/c/mlb-player-digital-engagement-forecasting/discussion/246499) who died in July 2019 still having targets increasing even in 2021. That had playerForTestSetAndFuturePreds  False so should not be included.\nBut what it showed was if players have other accounts like for Foundations, fund raising activities or are in the news for something like in that instance when staffers were charged and another on painkillers and his death, there were spikes in targets. None of that would be in the competition data. \n\nMy thinking is that the analytics side of this for the explainability prize is the only way to focus on those outside factors.  If there had been more clarity on use of external data and that data could be updated throughout the evaluation phase it might have helped.\n\nWhere the competition data fails also is in transaction dates that are after the fact or rosters are wrong, so players not on rosters are playing, trades, retires, call ups are announced and the spike shows the day previous. OK could be the rumour mill at work but just seems like the dates are wrong.\n\nAm waiting to see how the updated train to 20 Jul looks.",
    "1397048": "FYI - the promised \"opt-in\" 20 Jul data has been uploaded as `train_updated.csv` on the [data page](https://www.kaggle.com/c/mlb-player-digital-engagement-forecasting/data)"
  },
  "source": "meta"
}