{
  "id": 246499,
  "title": "Tyler Skaggs. What is the nature of the targets? ",
  "url": "/competitions/mlb-player-digital-engagement-forecasting/discussion/246499",
  "author_name": "",
  "post_date": "2021-06-15T14:33:48.914761500Z",
  "votes": 14,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Now, step by step I create <a href=\"https://www.kaggle.com/miklgr500/mlb-player-digital-engagement-forecasting-eda\" target=\"_blank\">EDA</a> research of MLB data and found a unique and interesting case of Taylor Skaggs(3.2.4).  It's case interesting that targets after the player dead increase. So it's a very strange situation if targets characterize player skills. So what is the real nature of targets? Are you have any idea?</p>",
  "messages": [
    {
      "id": "1350557",
      "postDate": "06/15/2021 14:33:48",
      "content": "<p>Now, step by step I create <a href=\"https://www.kaggle.com/miklgr500/mlb-player-digital-engagement-forecasting-eda\" target=\"_blank\">EDA</a> research of MLB data and found a unique and interesting case of Taylor Skaggs(3.2.4).  It's case interesting that targets after the player dead increase. So it's a very strange situation if targets characterize player skills. So what is the real nature of targets? Are you have any idea?</p>",
      "rawMarkdown": "Now, step by step I create [EDA](https://www.kaggle.com/miklgr500/mlb-player-digital-engagement-forecasting-eda) research of MLB data and found a unique and interesting case of Taylor Skaggs(3.2.4).  It's case interesting that targets after the player dead increase. So it's a very strange situation if targets characterize player skills. So what is the real nature of targets? Are you have any idea?",
      "votes": null
    },
    {
      "id": "1350675",
      "postDate": "06/15/2021 16:34:16",
      "content": "<p>The targets are (IMO) metrics derived from internet use.  As Kaggle is a Google company I would believe that they have a large number of metrics they track for advertising clicks, searches, etc.  </p>\n<p>The targets (once again IMO) do not characterize player skills but rather those skills drive us to have an interest in the player and engage in activities on the internet about that player.  </p>\n<p><a href=\"https://www.thinkwithgoogle.com/\" target=\"_blank\">thinkwithgoogle</a></p>",
      "rawMarkdown": "The targets are (IMO) metrics derived from internet use.  As Kaggle is a Google company I would believe that they have a large number of metrics they track for advertising clicks, searches, etc.  \n\nThe targets (once again IMO) do not characterize player skills but rather those skills drive us to have an interest in the player and engage in activities on the internet about that player.  \n\n[thinkwithgoogle](https://www.thinkwithgoogle.com/)",
      "votes": null
    },
    {
      "id": "1350781",
      "postDate": "06/15/2021 18:46:48",
      "content": "<p>Yes Kaggle is part of Google, but IMO, the four targets are the 4 twitter public_metrics retweet_count, reply_count, like_count, quote_count - this make sense as there are 4 public metrics in the Twitter API and they give us twitter followers.  I may be wrong, but I am working on comparing some player tweet stats with targets to see if I can find a link.</p>",
      "rawMarkdown": "Yes Kaggle is part of Google, but IMO, the four targets are the 4 twitter public_metrics retweet_count, reply_count, like_count, quote_count - this make sense as there are 4 public metrics in the Twitter API and they give us twitter followers.  I may be wrong, but I am working on comparing some player tweet stats with targets to see if I can find a link.",
      "votes": null
    },
    {
      "id": "1350789",
      "postDate": "06/15/2021 18:53:23",
      "content": "<p>As for the player issue you raise, my initial work with the data and engagement shows little correlation to players playing baseball and engagements seem to be tied way more to # of followers and player SM activity or a news story about the player…such as an untimely death.</p>",
      "rawMarkdown": "As for the player issue you raise, my initial work with the data and engagement shows little correlation to players playing baseball and engagements seem to be tied way more to # of followers and player SM activity or a news story about the player...such as an untimely death.",
      "votes": null
    },
    {
      "id": "1350794",
      "postDate": "06/15/2021 18:59:39",
      "content": "<p>If your correct than all players who do not have a twitter account/ followers should have zeros for all targets.  Likewise I think a team or two don't show twitter followers.  </p>\n<p>I have not looked that close - guess I need to- but I don't think there are many players with a set of target values at 0.0 for the entire 5 years.</p>\n<p>So I think your wrong - twitter is only a very small part of digital engagement.</p>\n<p>PS - still have my Orange from time I spent training for Six Sigma in Knoxville.</p>",
      "rawMarkdown": "If your correct than all players who do not have a twitter account/ followers should have zeros for all targets.  Likewise I think a team or two don't show twitter followers.  \n\nI have not looked that close - guess I need to- but I don't think there are many players with a set of target values at 0.0 for the entire 5 years.\n\nSo I think your wrong - twitter is only a very small part of digital engagement.\n\nPS - still have my Orange from time I spent training for Six Sigma in Knoxville.",
      "votes": null
    },
    {
      "id": "1350804",
      "postDate": "06/15/2021 19:18:26",
      "content": "<p>Agree - actual playing does not seem that important.  Found a blog with top endorsements (dollars) that seems a much nicer match to the data.   Getting news and twits into a model seems like it makes the most sense, but having that updated for the final scoring is way beyond my skill set.  Suppose one could work for winning an explainability prize - but…</p>",
      "rawMarkdown": "Agree - actual playing does not seem that important.  Found a blog with top endorsements (dollars) that seems a much nicer match to the data.   Getting news and twits into a model seems like it makes the most sense, but having that updated for the final scoring is way beyond my skill set.  Suppose one could work for winning an explainability prize - but...",
      "votes": null
    },
    {
      "id": "1350848",
      "postDate": "06/15/2021 20:36:32",
      "content": "<p>There is no internet access allowed for the competition which means getting updated news incorporated into the model won't be possible. Since there won't be updated external data during evaluation I think that a lot of the information from new data sources will already be incorporated into the average values of the target variables (though I think having access to the game schedule will likely be important). </p>",
      "rawMarkdown": "There is no internet access allowed for the competition which means getting updated news incorporated into the model won't be possible. Since there won't be updated external data during evaluation I think that a lot of the information from new data sources will already be incorporated into the average values of the target variables (though I think having access to the game schedule will likely be important).",
      "votes": null
    },
    {
      "id": "1354569",
      "postDate": "06/17/2021 17:04:13",
      "content": "<p>Has anyone looked at Google Trends for data on players? </p>",
      "rawMarkdown": "Has anyone looked at Google Trends for data on players?",
      "votes": null
    },
    {
      "id": "1354580",
      "postDate": "06/17/2021 17:19:02",
      "content": "<p><a href=\"https://www.kaggle.com/devinanzelmo\" target=\"_blank\">Devin's</a> comment  is the thing that keeps me focused on only the provided data right now - can you get enough score improvement with outside sources that you can survive the 45 days between the end of submissions and the final LB score.  </p>\n<p>Still looking to me that game and player stats are not going to be enough to completely predict engagement.  I think the recent internet buzz about pitchers using sticky stuff is not going to really show up in the stats enough to explain the engagement uptick I expect to see for those pitchers most likely to be sticky stuff users or those who \"confess\".</p>",
      "rawMarkdown": "[Devin's](https://www.kaggle.com/devinanzelmo) comment  is the thing that keeps me focused on only the provided data right now - can you get enough score improvement with outside sources that you can survive the 45 days between the end of submissions and the final LB score.  \n\nStill looking to me that game and player stats are not going to be enough to completely predict engagement.  I think the recent internet buzz about pitchers using sticky stuff is not going to really show up in the stats enough to explain the engagement uptick I expect to see for those pitchers most likely to be sticky stuff users or those who \"confess\".",
      "votes": null
    },
    {
      "id": "1356720",
      "postDate": "06/19/2021 07:22:14",
      "content": "<p>With respect to Tyler Skaggs (and potentially other players not in the current season), in players.csv there is a column playerForTestSetAndFuturePreds which is False in this case. I am assuming players for predictions will have True for now and the evaluation period and working with data for them.  Whether new players get added later, not sure, hopefully not.  But could be one of those unknowns to cater for and  has been mentioned elsewhere.  It is interesting though that there does not seem to be anything in transactions for this player like a Status Change or other type code if a player is no longer playing for whatever reason like injury or paternity leave, etc.  which do have transactions. </p>\n<p>Just looking at April 2021 for Tyler Skaggs, there are some increases in target4 which may coincide with the Tyler Skaggs Foundation - announcements on supporting local high school baseball in his memory, some fundraising at Angels games.<br>\nThere has also been news on court cases, charges, etc.  e.g. an LA Times article.  </p>\n<p>Whether this supports target4 relating to twitter activity, metrics or other trends not sure.  Theoretically updating a kaggle dataset that is publicly available (so adheres to external data rules should work) since notebooks are rerun during evaluation phase and would use the latest version of a dataset. But not sure that is the intention of this competition? </p>",
      "rawMarkdown": "With respect to Tyler Skaggs (and potentially other players not in the current season), in players.csv there is a column playerForTestSetAndFuturePreds which is False in this case. I am assuming players for predictions will have True for now and the evaluation period and working with data for them.  Whether new players get added later, not sure, hopefully not.  But could be one of those unknowns to cater for and  has been mentioned elsewhere.  It is interesting though that there does not seem to be anything in transactions for this player like a Status Change or other type code if a player is no longer playing for whatever reason like injury or paternity leave, etc.  which do have transactions. \n\nJust looking at April 2021 for Tyler Skaggs, there are some increases in target4 which may coincide with the Tyler Skaggs Foundation - announcements on supporting local high school baseball in his memory, some fundraising at Angels games.\nThere has also been news on court cases, charges, etc.  e.g. an LA Times article.  \n\nWhether this supports target4 relating to twitter activity, metrics or other trends not sure.  Theoretically updating a kaggle dataset that is publicly available (so adheres to external data rules should work) since notebooks are rerun during evaluation phase and would use the latest version of a dataset. But not sure that is the intention of this competition?",
      "votes": null
    },
    {
      "id": "1357366",
      "postDate": "06/19/2021 16:14:56",
      "content": "<p>Good point about updating a public dataset after the deadline for the competition. I am not familiar with how dataset/notebook versioning works. If notebooks automatically use the latest version of the dataset then this should work. </p>\n<p>I have a feeling that it is not the intention of the competition for this to be possible as there would be historical data available for every day but the last one in the final evaluation. This makes the all the effort they go through to have a good final test set sort of pointless. It is worth asking for a rules clarification. </p>",
      "rawMarkdown": "Good point about updating a public dataset after the deadline for the competition. I am not familiar with how dataset/notebook versioning works. If notebooks automatically use the latest version of the dataset then this should work. \n\nI have a feeling that it is not the intention of the competition for this to be possible as there would be historical data available for every day but the last one in the final evaluation. This makes the all the effort they go through to have a good final test set sort of pointless. It is worth asking for a rules clarification.",
      "votes": null
    },
    {
      "id": "1357381",
      "postDate": "06/19/2021 16:39:39",
      "content": "<p>In other discussion posts it's indicated that a new file with a new name will be added that contains data past the current 4/30 for the training set.  Kaggle does occasionally update a data set but this is normally the result of an error/leak.  I believe the SETI data is being/has been updated as a leak was discovered that permitted perfect scores.</p>\n<p>The only players that are being scored in the current public and future private test sets are those who show as True in 'playerForTestSetAndFuturePreds'.</p>\n<blockquote>\n  <p>playerForTestSetAndFuturePreds - Boolean, true if player is among those for whom predictions are to be made in test data&gt; </p>\n</blockquote>\n<p>I have seen nothing from the hosts that suggests more players will be added to the \"true\" list.  I would hope that they might drop a few players from this list if a players playing status changes after July 31 but not seen anything that would suggest that change either.</p>",
      "rawMarkdown": "In other discussion posts it's indicated that a new file with a new name will be added that contains data past the current 4/30 for the training set.  Kaggle does occasionally update a data set but this is normally the result of an error/leak.  I believe the SETI data is being/has been updated as a leak was discovered that permitted perfect scores.\n\nThe only players that are being scored in the current public and future private test sets are those who show as True in 'playerForTestSetAndFuturePreds'.\n> playerForTestSetAndFuturePreds - Boolean, true if player is among those for whom predictions are to be made in test data> \n\nI have seen nothing from the hosts that suggests more players will be added to the \"true\" list.  I would hope that they might drop a few players from this list if a players playing status changes after July 31 but not seen anything that would suggest that change either.",
      "votes": null
    },
    {
      "id": "1357998",
      "postDate": "06/20/2021 06:02:47",
      "content": "<p><a href=\"https://www.kaggle.com/devinanzelmo\" target=\"_blank\">@devinanzelmo</a> - for notebooks when they are rerun they should pick up the latest version of any data source used like output from another notebook or a kaggle dataset.  </p>\n<p>if what Ken Miller posted in here is correct - \"IMO, the four targets are the 4 twitter public_metrics retweet_count, reply_count, like_count, quote_count - this make sense as there are 4 public metrics in the Twitter API and they give us twitter followers\", <br>\nthese or other metrics may not be available until after the scoring for a particular day. but even having external data a day or two later might be useful because there seems to be a decay in targets after a big jump.  Of course, all this depends if correlations can be found for external metrics and targets that are better than competition data.   </p>",
      "rawMarkdown": "devinanzelmo - for notebooks when they are rerun they should pick up the latest version of any data source used like output from another notebook or a kaggle dataset.  \n    \nif what Ken Miller posted in here is correct - \"IMO, the four targets are the 4 twitter public_metrics retweet_count, reply_count, like_count, quote_count - this make sense as there are 4 public metrics in the Twitter API and they give us twitter followers\", \nthese or other metrics may not be available until after the scoring for a particular day. but even having external data a day or two later might be useful because there seems to be a decay in targets after a big jump.  Of course, all this depends if correlations can be found for external metrics and targets that are better than competition data.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1350675,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "06/15/2021 16:34:16",
      "content": "<p>The targets are (IMO) metrics derived from internet use.  As Kaggle is a Google company I would believe that they have a large number of metrics they track for advertising clicks, searches, etc.  </p>\n<p>The targets (once again IMO) do not characterize player skills but rather those skills drive us to have an interest in the player and engage in activities on the internet about that player.  </p>\n<p><a href=\"https://www.thinkwithgoogle.com/\" target=\"_blank\">thinkwithgoogle</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1350781,
      "author_name": "mlconsult",
      "author_url": "",
      "post_date": "06/15/2021 18:46:48",
      "content": "<p>Yes Kaggle is part of Google, but IMO, the four targets are the 4 twitter public_metrics retweet_count, reply_count, like_count, quote_count - this make sense as there are 4 public metrics in the Twitter API and they give us twitter followers.  I may be wrong, but I am working on comparing some player tweet stats with targets to see if I can find a link.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1350794,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "06/15/2021 18:59:39",
          "content": "<p>If your correct than all players who do not have a twitter account/ followers should have zeros for all targets.  Likewise I think a team or two don't show twitter followers.  </p>\n<p>I have not looked that close - guess I need to- but I don't think there are many players with a set of target values at 0.0 for the entire 5 years.</p>\n<p>So I think your wrong - twitter is only a very small part of digital engagement.</p>\n<p>PS - still have my Orange from time I spent training for Six Sigma in Knoxville.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1350789,
      "author_name": "mlconsult",
      "author_url": "",
      "post_date": "06/15/2021 18:53:23",
      "content": "<p>As for the player issue you raise, my initial work with the data and engagement shows little correlation to players playing baseball and engagements seem to be tied way more to # of followers and player SM activity or a news story about the player…such as an untimely death.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1350804,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "06/15/2021 19:18:26",
          "content": "<p>Agree - actual playing does not seem that important.  Found a blog with top endorsements (dollars) that seems a much nicer match to the data.   Getting news and twits into a model seems like it makes the most sense, but having that updated for the final scoring is way beyond my skill set.  Suppose one could work for winning an explainability prize - but…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1350848,
          "author_name": "devinanzelmo",
          "author_url": "",
          "post_date": "06/15/2021 20:36:32",
          "content": "<p>There is no internet access allowed for the competition which means getting updated news incorporated into the model won't be possible. Since there won't be updated external data during evaluation I think that a lot of the information from new data sources will already be incorporated into the average values of the target variables (though I think having access to the game schedule will likely be important). </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1354569,
          "author_name": "crained",
          "author_url": "",
          "post_date": "06/17/2021 17:04:13",
          "content": "<p>Has anyone looked at Google Trends for data on players? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1354580,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "06/17/2021 17:19:02",
          "content": "<p><a href=\"https://www.kaggle.com/devinanzelmo\" target=\"_blank\">Devin's</a> comment  is the thing that keeps me focused on only the provided data right now - can you get enough score improvement with outside sources that you can survive the 45 days between the end of submissions and the final LB score.  </p>\n<p>Still looking to me that game and player stats are not going to be enough to completely predict engagement.  I think the recent internet buzz about pitchers using sticky stuff is not going to really show up in the stats enough to explain the engagement uptick I expect to see for those pitchers most likely to be sticky stuff users or those who \"confess\".</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1356720,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "06/19/2021 07:22:14",
      "content": "<p>With respect to Tyler Skaggs (and potentially other players not in the current season), in players.csv there is a column playerForTestSetAndFuturePreds which is False in this case. I am assuming players for predictions will have True for now and the evaluation period and working with data for them.  Whether new players get added later, not sure, hopefully not.  But could be one of those unknowns to cater for and  has been mentioned elsewhere.  It is interesting though that there does not seem to be anything in transactions for this player like a Status Change or other type code if a player is no longer playing for whatever reason like injury or paternity leave, etc.  which do have transactions. </p>\n<p>Just looking at April 2021 for Tyler Skaggs, there are some increases in target4 which may coincide with the Tyler Skaggs Foundation - announcements on supporting local high school baseball in his memory, some fundraising at Angels games.<br>\nThere has also been news on court cases, charges, etc.  e.g. an LA Times article.  </p>\n<p>Whether this supports target4 relating to twitter activity, metrics or other trends not sure.  Theoretically updating a kaggle dataset that is publicly available (so adheres to external data rules should work) since notebooks are rerun during evaluation phase and would use the latest version of a dataset. But not sure that is the intention of this competition? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1357366,
          "author_name": "devinanzelmo",
          "author_url": "",
          "post_date": "06/19/2021 16:14:56",
          "content": "<p>Good point about updating a public dataset after the deadline for the competition. I am not familiar with how dataset/notebook versioning works. If notebooks automatically use the latest version of the dataset then this should work. </p>\n<p>I have a feeling that it is not the intention of the competition for this to be possible as there would be historical data available for every day but the last one in the final evaluation. This makes the all the effort they go through to have a good final test set sort of pointless. It is worth asking for a rules clarification. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1357381,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "06/19/2021 16:39:39",
          "content": "<p>In other discussion posts it's indicated that a new file with a new name will be added that contains data past the current 4/30 for the training set.  Kaggle does occasionally update a data set but this is normally the result of an error/leak.  I believe the SETI data is being/has been updated as a leak was discovered that permitted perfect scores.</p>\n<p>The only players that are being scored in the current public and future private test sets are those who show as True in 'playerForTestSetAndFuturePreds'.</p>\n<blockquote>\n  <p>playerForTestSetAndFuturePreds - Boolean, true if player is among those for whom predictions are to be made in test data&gt; </p>\n</blockquote>\n<p>I have seen nothing from the hosts that suggests more players will be added to the \"true\" list.  I would hope that they might drop a few players from this list if a players playing status changes after July 31 but not seen anything that would suggest that change either.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1357998,
          "author_name": "something4kag",
          "author_url": "",
          "post_date": "06/20/2021 06:02:47",
          "content": "<p><a href=\"https://www.kaggle.com/devinanzelmo\" target=\"_blank\">@devinanzelmo</a> - for notebooks when they are rerun they should pick up the latest version of any data source used like output from another notebook or a kaggle dataset.  </p>\n<p>if what Ken Miller posted in here is correct - \"IMO, the four targets are the 4 twitter public_metrics retweet_count, reply_count, like_count, quote_count - this make sense as there are 4 public metrics in the Twitter API and they give us twitter followers\", <br>\nthese or other metrics may not be available until after the scoring for a particular day. but even having external data a day or two later might be useful because there seems to be a decay in targets after a big jump.  Of course, all this depends if correlations can be found for external metrics and targets that are better than competition data.   </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1350557": "Now, step by step I create [EDA](https://www.kaggle.com/miklgr500/mlb-player-digital-engagement-forecasting-eda) research of MLB data and found a unique and interesting case of Taylor Skaggs(3.2.4).  It's case interesting that targets after the player dead increase. So it's a very strange situation if targets characterize player skills. So what is the real nature of targets? Are you have any idea?",
    "1350675": "The targets are (IMO) metrics derived from internet use.  As Kaggle is a Google company I would believe that they have a large number of metrics they track for advertising clicks, searches, etc.  \n\nThe targets (once again IMO) do not characterize player skills but rather those skills drive us to have an interest in the player and engage in activities on the internet about that player.  \n\n[thinkwithgoogle](https://www.thinkwithgoogle.com/)",
    "1350781": "Yes Kaggle is part of Google, but IMO, the four targets are the 4 twitter public_metrics retweet_count, reply_count, like_count, quote_count - this make sense as there are 4 public metrics in the Twitter API and they give us twitter followers.  I may be wrong, but I am working on comparing some player tweet stats with targets to see if I can find a link.",
    "1350789": "As for the player issue you raise, my initial work with the data and engagement shows little correlation to players playing baseball and engagements seem to be tied way more to # of followers and player SM activity or a news story about the player...such as an untimely death.",
    "1350794": "If your correct than all players who do not have a twitter account/ followers should have zeros for all targets.  Likewise I think a team or two don't show twitter followers.  \n\nI have not looked that close - guess I need to- but I don't think there are many players with a set of target values at 0.0 for the entire 5 years.\n\nSo I think your wrong - twitter is only a very small part of digital engagement.\n\nPS - still have my Orange from time I spent training for Six Sigma in Knoxville.",
    "1350804": "Agree - actual playing does not seem that important.  Found a blog with top endorsements (dollars) that seems a much nicer match to the data.   Getting news and twits into a model seems like it makes the most sense, but having that updated for the final scoring is way beyond my skill set.  Suppose one could work for winning an explainability prize - but...",
    "1350848": "There is no internet access allowed for the competition which means getting updated news incorporated into the model won't be possible. Since there won't be updated external data during evaluation I think that a lot of the information from new data sources will already be incorporated into the average values of the target variables (though I think having access to the game schedule will likely be important).",
    "1354569": "Has anyone looked at Google Trends for data on players?",
    "1354580": "[Devin's](https://www.kaggle.com/devinanzelmo) comment  is the thing that keeps me focused on only the provided data right now - can you get enough score improvement with outside sources that you can survive the 45 days between the end of submissions and the final LB score.  \n\nStill looking to me that game and player stats are not going to be enough to completely predict engagement.  I think the recent internet buzz about pitchers using sticky stuff is not going to really show up in the stats enough to explain the engagement uptick I expect to see for those pitchers most likely to be sticky stuff users or those who \"confess\".",
    "1356720": "With respect to Tyler Skaggs (and potentially other players not in the current season), in players.csv there is a column playerForTestSetAndFuturePreds which is False in this case. I am assuming players for predictions will have True for now and the evaluation period and working with data for them.  Whether new players get added later, not sure, hopefully not.  But could be one of those unknowns to cater for and  has been mentioned elsewhere.  It is interesting though that there does not seem to be anything in transactions for this player like a Status Change or other type code if a player is no longer playing for whatever reason like injury or paternity leave, etc.  which do have transactions. \n\nJust looking at April 2021 for Tyler Skaggs, there are some increases in target4 which may coincide with the Tyler Skaggs Foundation - announcements on supporting local high school baseball in his memory, some fundraising at Angels games.\nThere has also been news on court cases, charges, etc.  e.g. an LA Times article.  \n\nWhether this supports target4 relating to twitter activity, metrics or other trends not sure.  Theoretically updating a kaggle dataset that is publicly available (so adheres to external data rules should work) since notebooks are rerun during evaluation phase and would use the latest version of a dataset. But not sure that is the intention of this competition?",
    "1357366": "Good point about updating a public dataset after the deadline for the competition. I am not familiar with how dataset/notebook versioning works. If notebooks automatically use the latest version of the dataset then this should work. \n\nI have a feeling that it is not the intention of the competition for this to be possible as there would be historical data available for every day but the last one in the final evaluation. This makes the all the effort they go through to have a good final test set sort of pointless. It is worth asking for a rules clarification.",
    "1357381": "In other discussion posts it's indicated that a new file with a new name will be added that contains data past the current 4/30 for the training set.  Kaggle does occasionally update a data set but this is normally the result of an error/leak.  I believe the SETI data is being/has been updated as a leak was discovered that permitted perfect scores.\n\nThe only players that are being scored in the current public and future private test sets are those who show as True in 'playerForTestSetAndFuturePreds'.\n> playerForTestSetAndFuturePreds - Boolean, true if player is among those for whom predictions are to be made in test data> \n\nI have seen nothing from the hosts that suggests more players will be added to the \"true\" list.  I would hope that they might drop a few players from this list if a players playing status changes after July 31 but not seen anything that would suggest that change either.",
    "1357998": "devinanzelmo - for notebooks when they are rerun they should pick up the latest version of any data source used like output from another notebook or a kaggle dataset.  \n    \nif what Ken Miller posted in here is correct - \"IMO, the four targets are the 4 twitter public_metrics retweet_count, reply_count, like_count, quote_count - this make sense as there are 4 public metrics in the Twitter API and they give us twitter followers\", \nthese or other metrics may not be available until after the scoring for a particular day. but even having external data a day or two later might be useful because there seems to be a decay in targets after a big jump.  Of course, all this depends if correlations can be found for external metrics and targets that are better than competition data."
  },
  "source": "meta"
}