{
  "id": 275220,
  "title": "Papers on NFL and Machine Learning",
  "url": "/competitions/nfl-big-data-bowl-2022/discussion/275220",
  "author_name": "Gaju Ahmed",
  "post_date": "2021-09-29T12:00:29.711000",
  "votes": 5,
  "comment_count": 0,
  "views": 0,
  "content": "<p><a href=\"http://cs230.stanford.edu/projects_winter_2020/reports/32263160.pdf\"><strong>Deep Learning for In-Game NFL Predictions</strong> : </a> This project explores deep learning methods for predicting in-game NFL play outcomes. Systems that predict play outcomes in-game may help NFL team play-callers improve their strategies at a play-call time. I focus on two important outcomes, the yardage outcome of a play and the offensive play call (pass or run). My models combine pre-snap “situational data” with pre-snap images of the play. I run an array of learning models, including shallow CNNs and transfer learning, with and without the image data. The models show that while the features have no signal for yardage outcomes, offensive play calls can be predicted with much higher accuracy than benchmark models. However, the marginal value of the image data and sophisticated deep learning models is very low for both tasks.</p>\n<p><a href=\"https://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.1038.1982&amp;rep=rep1&amp;type=pdf\"><strong>A Hybrid Prediction System for American NFL Results</strong></a>This research work investigates the use of machine learning algorithms (Linear Regression and K-Nearest Neighbour) for NFL games result prediction. Data mining techniques were employed on carefully created features with datasets from NFL games statistics using RapidMiner and Java programming language in the backend. High attribute weights of features were obtained from the Linear Regression Model (LR) which provides a basis for the K-Nearest Neighbour Model (KNN). The result is a hybridized model which shows that using relevant features will provide good prediction accuracy. Unique features used are: Bookmakers betting spread and players’ performance metrics. The prediction accuracy of 80.65% obtained shows that the experiment is substantially better than many existing systems with accuracies of 59.4%, 60.7%, 65.05% and 67.08%. This can therefore be a reference point for future research in this area especially on employing machine learning in predictions.</p>\n<p><a href=\"https://dspace.mit.edu/handle/1721.1/129909\"> <strong>Leveraging machine learning to predict playcalling tendencies in the NFL</strong></a> In this thesis, we apply four machine learning models to NFL play-by-play data from 2009-2018 to predict whether a team will run or pass the ball on a given play. We tested our models using league-wide and team-specific data in five different situations on the field. Our best league-wide models achieved a test accuracy of 80% and our best team-specific models achieved a test accuracy of 86%. Relative to the baseline of the run-to-pass ratio, the best league-wide models achieved an increase in accuracy of 25% and the best team-specific models achieved an increase of 27%. Our models showed that the Tennessee Titans, the New York Jets, and the Cincinnati Bengals have been the most predictable offenses in the NFL over 10 years. We found that a team's in-game run-to-pass ratio and their win and score probabilities are the driving factors for offensive play-calling. Additionally, our results show that teams are more predictable later in games, and that less predictable teams tend to experience greater success offensively</p>\n<p><a href=\" https://arxiv.org/pdf/2109.08051.pdf\"><strong>Frame by frame completion probability of an NFL pass:</strong> </a> American football is an increasingly popular sport, with a growing audience in many countries in the world. The most-watched American football league in the world is the United States National Football League (NFL), where every offensive play can be either a run or a pass, and in this work, we focus on passes. Many factors can affect the probability of pass completion, such as receiver separation from the nearest defender, distance from the receiver to the passer, offense formation, among many others. When predicting the completion probability of a pass, it is essential to know who the target of the pass is. By using distance measures between players and the ball, it is possible to calculate empirical probabilities and predict very accurately who the target will be. The big question is: how likely is it for a pass to be completed in an NFL match while the ball is in the air? <strong>We developed a machine-learning algorithm to answer this based on several predictors.</strong> Using data from the 2018 NFL season, we obtained conditional and marginal predictions for pass completion probability based on a random forest model. This is based on a two-stage procedure: first, we calculate the probability of each offensive player being the pass target, then, conditional on the target, we predict completion probability based on the random forest model. Finally, the general completion probability can be calculated using the law of total probability. We present animations for selected plays and show the pass completion probability evolution.</p>\n<h4>About NFL Data</h4>\n<p><a href=\"https://etd.ohiolink.edu/apexprod/rws_etd/send_file/send?accession=bgsu1624982176445411&amp;disposition=inline\"><strong>MODERN ANALYSIS ON PASSING PLAYS IN THE NATIONAL FOOTBALL LEAGUE</strong>: </a>The National Football League is the most popular professional American Football League. The league publishes and sponsors a data competition on Kaggle called, “Big Data Bowl.” The inspiration for use of this dataset are commentators saying, “analytics does not account for everything.”</p>\n<p>I answered questions that are frequently debated on networks like ESPN. Those questions are, “does playing in a dome improve passing?”, “are better teams better at passing?”, and “are certain formations better at passing?” With the use of nonparametric statistical techniques, I can fnally give an answer with data evidence to these questions.</p>\n<p>This data set also has many variables, so we should consider reducing the dimension of the data. Through the use of Principal Component Analysis I can reduce the dimensions, while keeping interpretation and performance of the remaining data.</p>\n<p>Both nonparametric and multivariate analysis are based upon the on-site availability of the event. However, most features in American football datasets contain some type of time to event data. With the use of survival analysis I was able to examine whether different teams are better at completing passes.</p>\n<p>Finally, I want to discuss potential problems in my analysis. Daryl Morey, general manager for the Philadelphia 76ers, contends that “football is 10 years behind basketball and basketball is 10 years behind baseball.” Which I would like to expand on such as when we use response variables such as Expected Points Added</p>\n<p><a href=\"https://operations.nfl.com/gameday/analytics/stats-articles/\"><strong>THE EXTRA POINT</strong>: </a>Welcome to the Extra Point, where members of the NFL's football data and analytics team will share updates on league-wide trends in football data, interesting visualizations that showcase innovative ways to use the league's data and provide an inside look at how the NFL uses data-driven insight to improve and monitor player and team performance</p>",
  "messages": [
    {
      "id": 1528135,
      "postDate": "2021-09-29T12:00:29.710Z",
      "content": "<p><a href=\"http://cs230.stanford.edu/projects_winter_2020/reports/32263160.pdf\"><strong>Deep Learning for In-Game NFL Predictions</strong> : </a> This project explores deep learning methods for predicting in-game NFL play outcomes. Systems that predict play outcomes in-game may help NFL team play-callers improve their strategies at a play-call time. I focus on two important outcomes, the yardage outcome of a play and the offensive play call (pass or run). My models combine pre-snap “situational data” with pre-snap images of the play. I run an array of learning models, including shallow CNNs and transfer learning, with and without the image data. The models show that while the features have no signal for yardage outcomes, offensive play calls can be predicted with much higher accuracy than benchmark models. However, the marginal value of the image data and sophisticated deep learning models is very low for both tasks.</p>\n<p><a href=\"https://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.1038.1982&amp;rep=rep1&amp;type=pdf\"><strong>A Hybrid Prediction System for American NFL Results</strong></a>This research work investigates the use of machine learning algorithms (Linear Regression and K-Nearest Neighbour) for NFL games result prediction. Data mining techniques were employed on carefully created features with datasets from NFL games statistics using RapidMiner and Java programming language in the backend. High attribute weights of features were obtained from the Linear Regression Model (LR) which provides a basis for the K-Nearest Neighbour Model (KNN). The result is a hybridized model which shows that using relevant features will provide good prediction accuracy. Unique features used are: Bookmakers betting spread and players’ performance metrics. The prediction accuracy of 80.65% obtained shows that the experiment is substantially better than many existing systems with accuracies of 59.4%, 60.7%, 65.05% and 67.08%. This can therefore be a reference point for future research in this area especially on employing machine learning in predictions.</p>\n<p><a href=\"https://dspace.mit.edu/handle/1721.1/129909\"> <strong>Leveraging machine learning to predict playcalling tendencies in the NFL</strong></a> In this thesis, we apply four machine learning models to NFL play-by-play data from 2009-2018 to predict whether a team will run or pass the ball on a given play. We tested our models using league-wide and team-specific data in five different situations on the field. Our best league-wide models achieved a test accuracy of 80% and our best team-specific models achieved a test accuracy of 86%. Relative to the baseline of the run-to-pass ratio, the best league-wide models achieved an increase in accuracy of 25% and the best team-specific models achieved an increase of 27%. Our models showed that the Tennessee Titans, the New York Jets, and the Cincinnati Bengals have been the most predictable offenses in the NFL over 10 years. We found that a team's in-game run-to-pass ratio and their win and score probabilities are the driving factors for offensive play-calling. Additionally, our results show that teams are more predictable later in games, and that less predictable teams tend to experience greater success offensively</p>\n<p><a href=\" https://arxiv.org/pdf/2109.08051.pdf\"><strong>Frame by frame completion probability of an NFL pass:</strong> </a> American football is an increasingly popular sport, with a growing audience in many countries in the world. The most-watched American football league in the world is the United States National Football League (NFL), where every offensive play can be either a run or a pass, and in this work, we focus on passes. Many factors can affect the probability of pass completion, such as receiver separation from the nearest defender, distance from the receiver to the passer, offense formation, among many others. When predicting the completion probability of a pass, it is essential to know who the target of the pass is. By using distance measures between players and the ball, it is possible to calculate empirical probabilities and predict very accurately who the target will be. The big question is: how likely is it for a pass to be completed in an NFL match while the ball is in the air? <strong>We developed a machine-learning algorithm to answer this based on several predictors.</strong> Using data from the 2018 NFL season, we obtained conditional and marginal predictions for pass completion probability based on a random forest model. This is based on a two-stage procedure: first, we calculate the probability of each offensive player being the pass target, then, conditional on the target, we predict completion probability based on the random forest model. Finally, the general completion probability can be calculated using the law of total probability. We present animations for selected plays and show the pass completion probability evolution.</p>\n<h4>About NFL Data</h4>\n<p><a href=\"https://etd.ohiolink.edu/apexprod/rws_etd/send_file/send?accession=bgsu1624982176445411&amp;disposition=inline\"><strong>MODERN ANALYSIS ON PASSING PLAYS IN THE NATIONAL FOOTBALL LEAGUE</strong>: </a>The National Football League is the most popular professional American Football League. The league publishes and sponsors a data competition on Kaggle called, “Big Data Bowl.” The inspiration for use of this dataset are commentators saying, “analytics does not account for everything.”</p>\n<p>I answered questions that are frequently debated on networks like ESPN. Those questions are, “does playing in a dome improve passing?”, “are better teams better at passing?”, and “are certain formations better at passing?” With the use of nonparametric statistical techniques, I can fnally give an answer with data evidence to these questions.</p>\n<p>This data set also has many variables, so we should consider reducing the dimension of the data. Through the use of Principal Component Analysis I can reduce the dimensions, while keeping interpretation and performance of the remaining data.</p>\n<p>Both nonparametric and multivariate analysis are based upon the on-site availability of the event. However, most features in American football datasets contain some type of time to event data. With the use of survival analysis I was able to examine whether different teams are better at completing passes.</p>\n<p>Finally, I want to discuss potential problems in my analysis. Daryl Morey, general manager for the Philadelphia 76ers, contends that “football is 10 years behind basketball and basketball is 10 years behind baseball.” Which I would like to expand on such as when we use response variables such as Expected Points Added</p>\n<p><a href=\"https://operations.nfl.com/gameday/analytics/stats-articles/\"><strong>THE EXTRA POINT</strong>: </a>Welcome to the Extra Point, where members of the NFL's football data and analytics team will share updates on league-wide trends in football data, interesting visualizations that showcase innovative ways to use the league's data and provide an inside look at how the NFL uses data-driven insight to improve and monitor player and team performance</p>",
      "rawMarkdown": "<a href=\"http://cs230.stanford.edu/projects_winter_2020/reports/32263160.pdf\">**Deep Learning for In-Game NFL Predictions** : </a> This project explores deep learning methods for predicting in-game NFL play outcomes. Systems that predict play outcomes in-game may help NFL team play-callers improve their strategies at a play-call time. I focus on two important outcomes, the yardage outcome of a play and the offensive play call (pass or run). My models combine pre-snap “situational data” with pre-snap images of the play. I run an array of learning models, including shallow CNNs and transfer learning, with and without the image data. The models show that while the features have no signal for yardage outcomes, offensive play calls can be predicted with much higher accuracy than benchmark models. However, the marginal value of the image data and sophisticated deep learning models is very low for both tasks.\n\n<a href=\"https://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.1038.1982&rep=rep1&type=pdf\">**A Hybrid Prediction System for American NFL Results**</a>This research work investigates the use of machine learning algorithms (Linear Regression and K-Nearest Neighbour) for NFL games result prediction. Data mining techniques were employed on carefully created features with datasets from NFL games statistics using RapidMiner and Java programming language in the backend. High attribute weights of features were obtained from the Linear Regression Model (LR) which provides a basis for the K-Nearest Neighbour Model (KNN). The result is a hybridized model which shows that using relevant features will provide good prediction accuracy. Unique features used are: Bookmakers betting spread and players’ performance metrics. The prediction accuracy of 80.65% obtained shows that the experiment is substantially better than many existing systems with accuracies of 59.4%, 60.7%, 65.05% and 67.08%. This can therefore be a reference point for future research in this area especially on employing machine learning in predictions.\n\n<a href=\"https://dspace.mit.edu/handle/1721.1/129909\"> **Leveraging machine learning to predict playcalling tendencies in the NFL**</a> In this thesis, we apply four machine learning models to NFL play-by-play data from 2009-2018 to predict whether a team will run or pass the ball on a given play. We tested our models using league-wide and team-specific data in five different situations on the field. Our best league-wide models achieved a test accuracy of 80% and our best team-specific models achieved a test accuracy of 86%. Relative to the baseline of the run-to-pass ratio, the best league-wide models achieved an increase in accuracy of 25% and the best team-specific models achieved an increase of 27%. Our models showed that the Tennessee Titans, the New York Jets, and the Cincinnati Bengals have been the most predictable offenses in the NFL over 10 years. We found that a team's in-game run-to-pass ratio and their win and score probabilities are the driving factors for offensive play-calling. Additionally, our results show that teams are more predictable later in games, and that less predictable teams tend to experience greater success offensively\n\n<a href=\" https://arxiv.org/pdf/2109.08051.pdf\">**Frame by frame completion probability of an NFL pass:** </a> American football is an increasingly popular sport, with a growing audience in many countries in the world. The most-watched American football league in the world is the United States National Football League (NFL), where every offensive play can be either a run or a pass, and in this work, we focus on passes. Many factors can affect the probability of pass completion, such as receiver separation from the nearest defender, distance from the receiver to the passer, offense formation, among many others. When predicting the completion probability of a pass, it is essential to know who the target of the pass is. By using distance measures between players and the ball, it is possible to calculate empirical probabilities and predict very accurately who the target will be. The big question is: how likely is it for a pass to be completed in an NFL match while the ball is in the air? **We developed a machine-learning algorithm to answer this based on several predictors.** Using data from the 2018 NFL season, we obtained conditional and marginal predictions for pass completion probability based on a random forest model. This is based on a two-stage procedure: first, we calculate the probability of each offensive player being the pass target, then, conditional on the target, we predict completion probability based on the random forest model. Finally, the general completion probability can be calculated using the law of total probability. We present animations for selected plays and show the pass completion probability evolution.\n\n\n#### About NFL Data\n<a href=\"https://etd.ohiolink.edu/apexprod/rws_etd/send_file/send?accession=bgsu1624982176445411&disposition=inline\">**MODERN ANALYSIS ON PASSING PLAYS IN THE NATIONAL FOOTBALL LEAGUE**: </a>The National Football League is the most popular professional American Football League. The league publishes and sponsors a data competition on Kaggle called, “Big Data Bowl.” The inspiration for use of this dataset are commentators saying, “analytics does not account for everything.”\n\nI answered questions that are frequently debated on networks like ESPN. Those questions are, “does playing in a dome improve passing?”, “are better teams better at passing?”, and “are certain formations better at passing?” With the use of nonparametric statistical techniques, I can fnally give an answer with data evidence to these questions.\n\nThis data set also has many variables, so we should consider reducing the dimension of the data. Through the use of Principal Component Analysis I can reduce the dimensions, while keeping interpretation and performance of the remaining data.\n\nBoth nonparametric and multivariate analysis are based upon the on-site availability of the event. However, most features in American football datasets contain some type of time to event data. With the use of survival analysis I was able to examine whether different teams are better at completing passes.\n\nFinally, I want to discuss potential problems in my analysis. Daryl Morey, general manager for the Philadelphia 76ers, contends that “football is 10 years behind basketball and basketball is 10 years behind baseball.” Which I would like to expand on such as when we use response variables such as Expected Points Added\n\n<a href = \"https://operations.nfl.com/gameday/analytics/stats-articles/\">**THE EXTRA POINT**: </a>Welcome to the Extra Point, where members of the NFL's football data and analytics team will share updates on league-wide trends in football data, interesting visualizations that showcase innovative ways to use the league's data and provide an inside look at how the NFL uses data-driven insight to improve and monitor player and team performance",
      "votes": 5
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1528135": "<a href=\"http://cs230.stanford.edu/projects_winter_2020/reports/32263160.pdf\">**Deep Learning for In-Game NFL Predictions** : </a> This project explores deep learning methods for predicting in-game NFL play outcomes. Systems that predict play outcomes in-game may help NFL team play-callers improve their strategies at a play-call time. I focus on two important outcomes, the yardage outcome of a play and the offensive play call (pass or run). My models combine pre-snap “situational data” with pre-snap images of the play. I run an array of learning models, including shallow CNNs and transfer learning, with and without the image data. The models show that while the features have no signal for yardage outcomes, offensive play calls can be predicted with much higher accuracy than benchmark models. However, the marginal value of the image data and sophisticated deep learning models is very low for both tasks.\n\n<a href=\"https://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.1038.1982&rep=rep1&type=pdf\">**A Hybrid Prediction System for American NFL Results**</a>This research work investigates the use of machine learning algorithms (Linear Regression and K-Nearest Neighbour) for NFL games result prediction. Data mining techniques were employed on carefully created features with datasets from NFL games statistics using RapidMiner and Java programming language in the backend. High attribute weights of features were obtained from the Linear Regression Model (LR) which provides a basis for the K-Nearest Neighbour Model (KNN). The result is a hybridized model which shows that using relevant features will provide good prediction accuracy. Unique features used are: Bookmakers betting spread and players’ performance metrics. The prediction accuracy of 80.65% obtained shows that the experiment is substantially better than many existing systems with accuracies of 59.4%, 60.7%, 65.05% and 67.08%. This can therefore be a reference point for future research in this area especially on employing machine learning in predictions.\n\n<a href=\"https://dspace.mit.edu/handle/1721.1/129909\"> **Leveraging machine learning to predict playcalling tendencies in the NFL**</a> In this thesis, we apply four machine learning models to NFL play-by-play data from 2009-2018 to predict whether a team will run or pass the ball on a given play. We tested our models using league-wide and team-specific data in five different situations on the field. Our best league-wide models achieved a test accuracy of 80% and our best team-specific models achieved a test accuracy of 86%. Relative to the baseline of the run-to-pass ratio, the best league-wide models achieved an increase in accuracy of 25% and the best team-specific models achieved an increase of 27%. Our models showed that the Tennessee Titans, the New York Jets, and the Cincinnati Bengals have been the most predictable offenses in the NFL over 10 years. We found that a team's in-game run-to-pass ratio and their win and score probabilities are the driving factors for offensive play-calling. Additionally, our results show that teams are more predictable later in games, and that less predictable teams tend to experience greater success offensively\n\n<a href=\" https://arxiv.org/pdf/2109.08051.pdf\">**Frame by frame completion probability of an NFL pass:** </a> American football is an increasingly popular sport, with a growing audience in many countries in the world. The most-watched American football league in the world is the United States National Football League (NFL), where every offensive play can be either a run or a pass, and in this work, we focus on passes. Many factors can affect the probability of pass completion, such as receiver separation from the nearest defender, distance from the receiver to the passer, offense formation, among many others. When predicting the completion probability of a pass, it is essential to know who the target of the pass is. By using distance measures between players and the ball, it is possible to calculate empirical probabilities and predict very accurately who the target will be. The big question is: how likely is it for a pass to be completed in an NFL match while the ball is in the air? **We developed a machine-learning algorithm to answer this based on several predictors.** Using data from the 2018 NFL season, we obtained conditional and marginal predictions for pass completion probability based on a random forest model. This is based on a two-stage procedure: first, we calculate the probability of each offensive player being the pass target, then, conditional on the target, we predict completion probability based on the random forest model. Finally, the general completion probability can be calculated using the law of total probability. We present animations for selected plays and show the pass completion probability evolution.\n\n\n#### About NFL Data\n<a href=\"https://etd.ohiolink.edu/apexprod/rws_etd/send_file/send?accession=bgsu1624982176445411&disposition=inline\">**MODERN ANALYSIS ON PASSING PLAYS IN THE NATIONAL FOOTBALL LEAGUE**: </a>The National Football League is the most popular professional American Football League. The league publishes and sponsors a data competition on Kaggle called, “Big Data Bowl.” The inspiration for use of this dataset are commentators saying, “analytics does not account for everything.”\n\nI answered questions that are frequently debated on networks like ESPN. Those questions are, “does playing in a dome improve passing?”, “are better teams better at passing?”, and “are certain formations better at passing?” With the use of nonparametric statistical techniques, I can fnally give an answer with data evidence to these questions.\n\nThis data set also has many variables, so we should consider reducing the dimension of the data. Through the use of Principal Component Analysis I can reduce the dimensions, while keeping interpretation and performance of the remaining data.\n\nBoth nonparametric and multivariate analysis are based upon the on-site availability of the event. However, most features in American football datasets contain some type of time to event data. With the use of survival analysis I was able to examine whether different teams are better at completing passes.\n\nFinally, I want to discuss potential problems in my analysis. Daryl Morey, general manager for the Philadelphia 76ers, contends that “football is 10 years behind basketball and basketball is 10 years behind baseball.” Which I would like to expand on such as when we use response variables such as Expected Points Added\n\n<a href = \"https://operations.nfl.com/gameday/analytics/stats-articles/\">**THE EXTRA POINT**: </a>Welcome to the Extra Point, where members of the NFL's football data and analytics team will share updates on league-wide trends in football data, interesting visualizations that showcase innovative ways to use the league's data and provide an inside look at how the NFL uses data-driven insight to improve and monitor player and team performance"
  }
}