{
  "id": 245382,
  "title": "Papers on Baseball and Machine Learning",
  "url": "/competitions/mlb-player-digital-engagement-forecasting/discussion/245382",
  "author_name": "Charlie Craine",
  "post_date": "2021-06-10T18:20:02.684000",
  "votes": 13,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hey everyone!</p>\n<p>I wanted to start a papers thread and build on it, and hope others share as well, papers to gain domain knowledge. I have some domain knowledge in this area so I have downloaded papers that appeared to be relevant after reading their abstracts. I'll be reading these over the coming days and commenting more as I go through all of them. Hope this helps!</p>\n<p><strong>Research Papers:</strong></p>\n<ul>\n<li><p><a href=\"https://arxiv.org/abs/1811.07259\" target=\"_blank\">Modeling Baseball Outcomes as Higher-Order Markov Chains</a> - This paper deals with modeling the match outcomes of each of the ten teams in the KBO league as a higher-order Markov chain, where the possible states are win (\"W\"), draw (\"D\"), and loss (\"L\"). For each team, the value of k in which the kth order Markov chain model best describes the match outcome sequence is computed. Further, whether there are any patterns between such a value of k and the team's overall performance in the league is examined. We find that for the top three teams in the league, lower values of k tend to have the kth order Markov chain to better model their outcome, but the other teams don't reveal such patterns.</p></li>\n<li><p><a href=\"https://arxiv.org/abs/2103.12094\" target=\"_blank\">Modelling intransitivity in pairwise comparisons with application to baseball data</a> - In most commonly used ranking systems, some level of underlying transitivity is assumed. If transitivity exists in a system then information about pairwise comparisons can be translated to other linked pairs. For example, if typically A beats B and B beats C, this could inform us about the expected outcome between A and C. We show that in the seminal Bradley-Terry model knowing the probabilities of A beating B and B beating C completely defines the probability of A beating C, with these probabilities determined by individual skill levels of A, B and C.</p></li>\n<li><p><a href=\"https://arxiv.org/abs/1712.05754\" target=\"_blank\">Understanding Career Progression in Baseball Through Machine Learning</a> - Professional baseball players are increasingly guaranteed expensive long-term contracts, with over 70 deals signed in excess of $90 million, mostly in the last decade. These are substantial sums compared to a typical franchise valuation of $1-2 billion. Hence, the players to whom a team chooses to give such a contract can have an enormous impact on both competitiveness and profit. Despite this, most published approaches examining career progression in baseball are fairly simplistic. We applied four machine learning algorithms to the problem and soundly improved upon existing approaches, particularly for batting data.</p></li>\n<li><p><a href=\"https://arxiv.org/abs/2006.03348\" target=\"_blank\">Same-Score Streaks: A Case Study in Probability Modeling</a> - A same-score streak in sports is a sequence of games where the scores are equivalent in all games. The motivating problem arose from college basketball, however due to the difficulty in collecting data, streaks in Major League Baseball (MLB) were studied instead. This paper explores the historic data from regular-season games between 1901 and 2019 to include the likelihood of streaks of length 2, 3 and 4. Then we explore various probability models for the distribution of runs scored during MLB games and seasons and generate simulated statistics for the same length of streaks.</p></li>\n<li><p><a href=\"https://arxiv.org/abs/2009.00550\" target=\"_blank\">Using Social Networks to Improve Group Transition Prediction in Professional Sports</a> - We examine whether social data can be used to predict how members of Major League Baseball (MLB) and members of the National Basketball Association (NBA) transition between teams during their career. We find that incorporating social data into various machine learning algorithms substantially improves the algorithms' ability to correctly determine these transitions. In particular, we measure how player performance, team fitness, and social data individually and collectively contribute to predicting these transitions. Incorporating individual performance and team fitness both improve the predictive accuracy of our algorithms. However, this improvement is dwarfed by the improvement seen when we include social data suggesting that social relationships have a comparatively large effect on player transitions in both MLB and in the NBA.</p></li>\n</ul>\n<p><strong>Bonus domain:</strong></p>\n<ul>\n<li><p><a href=\"https://arxiv.org/abs/1906.05029\" target=\"_blank\">Who Will Win It? An In-game Win Probability Model for Football</a> - In-game win probability is a statistical metric that provides a sports team's likelihood of winning at any given point in a game, based on the performance of historical teams in the same situation. In-game win-probability models have been extensively studied in baseball, basketball and American football. </p></li>\n<li><p><a href=\"https://arxiv.org/abs/2011.09192\" target=\"_blank\">Game Plan: What AI can do for Football, and What Football can do for AI</a> - The rapid progress in artificial intelligence (AI) and machine learning has opened unprecedented analytics possibilities in various team and individual sports, including baseball, basketball, and tennis. More recently, AI techniques have been applied to football, due to a huge increase in data collection by professional teams, increased computational power, and advances in machine learning, with the goal of better addressing new scientific challenges involved in the analysis of both individual players' and coordinated teams' behaviors.</p></li>\n</ul>",
  "messages": [
    {
      "id": 1344264,
      "postDate": "2021-06-10T18:20:02.683Z",
      "content": "<p>Hey everyone!</p>\n<p>I wanted to start a papers thread and build on it, and hope others share as well, papers to gain domain knowledge. I have some domain knowledge in this area so I have downloaded papers that appeared to be relevant after reading their abstracts. I'll be reading these over the coming days and commenting more as I go through all of them. Hope this helps!</p>\n<p><strong>Research Papers:</strong></p>\n<ul>\n<li><p><a href=\"https://arxiv.org/abs/1811.07259\" target=\"_blank\">Modeling Baseball Outcomes as Higher-Order Markov Chains</a> - This paper deals with modeling the match outcomes of each of the ten teams in the KBO league as a higher-order Markov chain, where the possible states are win (\"W\"), draw (\"D\"), and loss (\"L\"). For each team, the value of k in which the kth order Markov chain model best describes the match outcome sequence is computed. Further, whether there are any patterns between such a value of k and the team's overall performance in the league is examined. We find that for the top three teams in the league, lower values of k tend to have the kth order Markov chain to better model their outcome, but the other teams don't reveal such patterns.</p></li>\n<li><p><a href=\"https://arxiv.org/abs/2103.12094\" target=\"_blank\">Modelling intransitivity in pairwise comparisons with application to baseball data</a> - In most commonly used ranking systems, some level of underlying transitivity is assumed. If transitivity exists in a system then information about pairwise comparisons can be translated to other linked pairs. For example, if typically A beats B and B beats C, this could inform us about the expected outcome between A and C. We show that in the seminal Bradley-Terry model knowing the probabilities of A beating B and B beating C completely defines the probability of A beating C, with these probabilities determined by individual skill levels of A, B and C.</p></li>\n<li><p><a href=\"https://arxiv.org/abs/1712.05754\" target=\"_blank\">Understanding Career Progression in Baseball Through Machine Learning</a> - Professional baseball players are increasingly guaranteed expensive long-term contracts, with over 70 deals signed in excess of $90 million, mostly in the last decade. These are substantial sums compared to a typical franchise valuation of $1-2 billion. Hence, the players to whom a team chooses to give such a contract can have an enormous impact on both competitiveness and profit. Despite this, most published approaches examining career progression in baseball are fairly simplistic. We applied four machine learning algorithms to the problem and soundly improved upon existing approaches, particularly for batting data.</p></li>\n<li><p><a href=\"https://arxiv.org/abs/2006.03348\" target=\"_blank\">Same-Score Streaks: A Case Study in Probability Modeling</a> - A same-score streak in sports is a sequence of games where the scores are equivalent in all games. The motivating problem arose from college basketball, however due to the difficulty in collecting data, streaks in Major League Baseball (MLB) were studied instead. This paper explores the historic data from regular-season games between 1901 and 2019 to include the likelihood of streaks of length 2, 3 and 4. Then we explore various probability models for the distribution of runs scored during MLB games and seasons and generate simulated statistics for the same length of streaks.</p></li>\n<li><p><a href=\"https://arxiv.org/abs/2009.00550\" target=\"_blank\">Using Social Networks to Improve Group Transition Prediction in Professional Sports</a> - We examine whether social data can be used to predict how members of Major League Baseball (MLB) and members of the National Basketball Association (NBA) transition between teams during their career. We find that incorporating social data into various machine learning algorithms substantially improves the algorithms' ability to correctly determine these transitions. In particular, we measure how player performance, team fitness, and social data individually and collectively contribute to predicting these transitions. Incorporating individual performance and team fitness both improve the predictive accuracy of our algorithms. However, this improvement is dwarfed by the improvement seen when we include social data suggesting that social relationships have a comparatively large effect on player transitions in both MLB and in the NBA.</p></li>\n</ul>\n<p><strong>Bonus domain:</strong></p>\n<ul>\n<li><p><a href=\"https://arxiv.org/abs/1906.05029\" target=\"_blank\">Who Will Win It? An In-game Win Probability Model for Football</a> - In-game win probability is a statistical metric that provides a sports team's likelihood of winning at any given point in a game, based on the performance of historical teams in the same situation. In-game win-probability models have been extensively studied in baseball, basketball and American football. </p></li>\n<li><p><a href=\"https://arxiv.org/abs/2011.09192\" target=\"_blank\">Game Plan: What AI can do for Football, and What Football can do for AI</a> - The rapid progress in artificial intelligence (AI) and machine learning has opened unprecedented analytics possibilities in various team and individual sports, including baseball, basketball, and tennis. More recently, AI techniques have been applied to football, due to a huge increase in data collection by professional teams, increased computational power, and advances in machine learning, with the goal of better addressing new scientific challenges involved in the analysis of both individual players' and coordinated teams' behaviors.</p></li>\n</ul>",
      "rawMarkdown": "Hey everyone!\n\nI wanted to start a papers thread and build on it, and hope others share as well, papers to gain domain knowledge. I have some domain knowledge in this area so I have downloaded papers that appeared to be relevant after reading their abstracts. I'll be reading these over the coming days and commenting more as I go through all of them. Hope this helps!\n\n**Research Papers:**\n\n- [Modeling Baseball Outcomes as Higher-Order Markov Chains](https://arxiv.org/abs/1811.07259) - This paper deals with modeling the match outcomes of each of the ten teams in the KBO league as a higher-order Markov chain, where the possible states are win (\"W\"), draw (\"D\"), and loss (\"L\"). For each team, the value of k in which the kth order Markov chain model best describes the match outcome sequence is computed. Further, whether there are any patterns between such a value of k and the team's overall performance in the league is examined. We find that for the top three teams in the league, lower values of k tend to have the kth order Markov chain to better model their outcome, but the other teams don't reveal such patterns.\n\n- [Modelling intransitivity in pairwise comparisons with application to baseball data](https://arxiv.org/abs/2103.12094) - In most commonly used ranking systems, some level of underlying transitivity is assumed. If transitivity exists in a system then information about pairwise comparisons can be translated to other linked pairs. For example, if typically A beats B and B beats C, this could inform us about the expected outcome between A and C. We show that in the seminal Bradley-Terry model knowing the probabilities of A beating B and B beating C completely defines the probability of A beating C, with these probabilities determined by individual skill levels of A, B and C.\n\n- [Understanding Career Progression in Baseball Through Machine Learning](https://arxiv.org/abs/1712.05754) - Professional baseball players are increasingly guaranteed expensive long-term contracts, with over 70 deals signed in excess of $90 million, mostly in the last decade. These are substantial sums compared to a typical franchise valuation of $1-2 billion. Hence, the players to whom a team chooses to give such a contract can have an enormous impact on both competitiveness and profit. Despite this, most published approaches examining career progression in baseball are fairly simplistic. We applied four machine learning algorithms to the problem and soundly improved upon existing approaches, particularly for batting data.\n\n- [Same-Score Streaks: A Case Study in Probability Modeling](https://arxiv.org/abs/2006.03348) - A same-score streak in sports is a sequence of games where the scores are equivalent in all games. The motivating problem arose from college basketball, however due to the difficulty in collecting data, streaks in Major League Baseball (MLB) were studied instead. This paper explores the historic data from regular-season games between 1901 and 2019 to include the likelihood of streaks of length 2, 3 and 4. Then we explore various probability models for the distribution of runs scored during MLB games and seasons and generate simulated statistics for the same length of streaks.\n\n- [Using Social Networks to Improve Group Transition Prediction in Professional Sports](https://arxiv.org/abs/2009.00550) - We examine whether social data can be used to predict how members of Major League Baseball (MLB) and members of the National Basketball Association (NBA) transition between teams during their career. We find that incorporating social data into various machine learning algorithms substantially improves the algorithms' ability to correctly determine these transitions. In particular, we measure how player performance, team fitness, and social data individually and collectively contribute to predicting these transitions. Incorporating individual performance and team fitness both improve the predictive accuracy of our algorithms. However, this improvement is dwarfed by the improvement seen when we include social data suggesting that social relationships have a comparatively large effect on player transitions in both MLB and in the NBA.\n\n**Bonus domain:**\n- [Who Will Win It? An In-game Win Probability Model for Football](https://arxiv.org/abs/1906.05029) - In-game win probability is a statistical metric that provides a sports team's likelihood of winning at any given point in a game, based on the performance of historical teams in the same situation. In-game win-probability models have been extensively studied in baseball, basketball and American football. \n\n- [Game Plan: What AI can do for Football, and What Football can do for AI](https://arxiv.org/abs/2011.09192) - The rapid progress in artificial intelligence (AI) and machine learning has opened unprecedented analytics possibilities in various team and individual sports, including baseball, basketball, and tennis. More recently, AI techniques have been applied to football, due to a huge increase in data collection by professional teams, increased computational power, and advances in machine learning, with the goal of better addressing new scientific challenges involved in the analysis of both individual players' and coordinated teams' behaviors.",
      "votes": 13
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1344264": "Hey everyone!\n\nI wanted to start a papers thread and build on it, and hope others share as well, papers to gain domain knowledge. I have some domain knowledge in this area so I have downloaded papers that appeared to be relevant after reading their abstracts. I'll be reading these over the coming days and commenting more as I go through all of them. Hope this helps!\n\n**Research Papers:**\n\n- [Modeling Baseball Outcomes as Higher-Order Markov Chains](https://arxiv.org/abs/1811.07259) - This paper deals with modeling the match outcomes of each of the ten teams in the KBO league as a higher-order Markov chain, where the possible states are win (\"W\"), draw (\"D\"), and loss (\"L\"). For each team, the value of k in which the kth order Markov chain model best describes the match outcome sequence is computed. Further, whether there are any patterns between such a value of k and the team's overall performance in the league is examined. We find that for the top three teams in the league, lower values of k tend to have the kth order Markov chain to better model their outcome, but the other teams don't reveal such patterns.\n\n- [Modelling intransitivity in pairwise comparisons with application to baseball data](https://arxiv.org/abs/2103.12094) - In most commonly used ranking systems, some level of underlying transitivity is assumed. If transitivity exists in a system then information about pairwise comparisons can be translated to other linked pairs. For example, if typically A beats B and B beats C, this could inform us about the expected outcome between A and C. We show that in the seminal Bradley-Terry model knowing the probabilities of A beating B and B beating C completely defines the probability of A beating C, with these probabilities determined by individual skill levels of A, B and C.\n\n- [Understanding Career Progression in Baseball Through Machine Learning](https://arxiv.org/abs/1712.05754) - Professional baseball players are increasingly guaranteed expensive long-term contracts, with over 70 deals signed in excess of $90 million, mostly in the last decade. These are substantial sums compared to a typical franchise valuation of $1-2 billion. Hence, the players to whom a team chooses to give such a contract can have an enormous impact on both competitiveness and profit. Despite this, most published approaches examining career progression in baseball are fairly simplistic. We applied four machine learning algorithms to the problem and soundly improved upon existing approaches, particularly for batting data.\n\n- [Same-Score Streaks: A Case Study in Probability Modeling](https://arxiv.org/abs/2006.03348) - A same-score streak in sports is a sequence of games where the scores are equivalent in all games. The motivating problem arose from college basketball, however due to the difficulty in collecting data, streaks in Major League Baseball (MLB) were studied instead. This paper explores the historic data from regular-season games between 1901 and 2019 to include the likelihood of streaks of length 2, 3 and 4. Then we explore various probability models for the distribution of runs scored during MLB games and seasons and generate simulated statistics for the same length of streaks.\n\n- [Using Social Networks to Improve Group Transition Prediction in Professional Sports](https://arxiv.org/abs/2009.00550) - We examine whether social data can be used to predict how members of Major League Baseball (MLB) and members of the National Basketball Association (NBA) transition between teams during their career. We find that incorporating social data into various machine learning algorithms substantially improves the algorithms' ability to correctly determine these transitions. In particular, we measure how player performance, team fitness, and social data individually and collectively contribute to predicting these transitions. Incorporating individual performance and team fitness both improve the predictive accuracy of our algorithms. However, this improvement is dwarfed by the improvement seen when we include social data suggesting that social relationships have a comparatively large effect on player transitions in both MLB and in the NBA.\n\n**Bonus domain:**\n- [Who Will Win It? An In-game Win Probability Model for Football](https://arxiv.org/abs/1906.05029) - In-game win probability is a statistical metric that provides a sports team's likelihood of winning at any given point in a game, based on the performance of historical teams in the same situation. In-game win-probability models have been extensively studied in baseball, basketball and American football. \n\n- [Game Plan: What AI can do for Football, and What Football can do for AI](https://arxiv.org/abs/2011.09192) - The rapid progress in artificial intelligence (AI) and machine learning has opened unprecedented analytics possibilities in various team and individual sports, including baseball, basketball, and tennis. More recently, AI techniques have been applied to football, due to a huge increase in data collection by professional teams, increased computational power, and advances in machine learning, with the goal of better addressing new scientific challenges involved in the analysis of both individual players' and coordinated teams' behaviors."
  }
}