{
  "id": 152213,
  "title": "Winning solution recap - A score for competitiveness",
  "url": "/competitions/march-madness-analytics-2020/discussion/152213",
  "author_name": "Luca Basanisi",
  "post_date": "2020-05-18T23:16:17.724000",
  "votes": 6,
  "comment_count": 0,
  "views": 0,
  "content": "<p>This is the first time I win a competition here on Kaggle and I was definitely not expecting to end up in the top 5, so I will try my best in further explaining my solution.</p>\n\n<p>I want to start by thanking the Kaggle team for organizing the competition and providing good quality data for us to have fun with. This is by far my favorite recurring competition, despite the fact that this year's circumstances deprived us of the thrill of March Madness.</p>\n\n<p>Secondly, I would like to thank all the competitors and content creators of all 3 competitions (the two that got canceled and this one) because they have been a source of inspiration to get better.</p>\n\n<p>Here you can find my solution <a href=\"https://www.kaggle.com/lucabasa/quantify-the-madness-a-study-of-competitiveness\">https://www.kaggle.com/lucabasa/quantify-the-madness-a-study-of-competitiveness</a></p>\n\n<h1>The idea behind it</h1>\n\n<p>I took inspiration from the competition description and I spent quite some time thinking about what competitiveness would mean. The main drive for me has been exploring the play by play data, as I never used them in previous editions and I was curious to see what can I with them. I thought they were more useful than the simple game stats as it can be that stats do not tell the full story.</p>\n\n<p>I started simple and guessed that a competitive game should not finish with a large point difference, which is kind of obvious. The second step was to recognize that sometimes a team just catches fire and blows the other team away in the last part of the game. Therefore, I decided to include other criteria to account for lead changes and point differences at different stages of the game. In the notebook you will find the details and in the utility scripts you can find how they were computed.</p>\n\n<p>I then chose 6 criteria and arbitrarily decided that I was calling those games competitive. The decision was driven by the distribution of these characteristics (like the point difference at the 37th minute of the game) and get something that makes sense basketball-wise. At this point, each criteria was selecting about 10-20% of the games as competitive.</p>\n\n<h1>The model</h1>\n\n<p>The idea is to predict if the hand labeling done above is something predictable by a machine learning model, given the characteristics of the game. Given how the data is provided, I made some manipulation that can be summarized as follows:\n* Getting the total number of each statistic. For example the total number of rebounds\n* Getting the difference of each statistic. Like the difference in Steals. This is given in absolute values as I did not want the model to use who won. An alternative approach would be to duplicate the dataset (and inverting winning and losing team) as the popular raddar approach suggests (not sure @raddar is the original author of it but those notebooks are the first one I saw a couple of years ago)\n* Doing the above 2 points for different stages of the game. For example at half time, at the 37th minute, and at the end of the game.</p>\n\n<p>I wanted the model to be fairly simple as the story would have been about explaining it. At the same time, I wanted to practice a few things, namely getting partial dependence plot in a kfold validation scheme. I adopted a 5 fold validation and made sure that there was no obvious leakage of the target into the training set.</p>\n\n<p>Most of the modeling decisions were made by iterating over a Logistic Regression model. This was done because LR is simple to understand and blazing fast. At the end XGBoost showed the best performance, I don't exclude LightGBM would do better but I don't know how to tune it properly.</p>\n\n<p>The final decision was made by manually inspecting a few misclassified samples. In particular, with my great surprise, they were mostly games that were mislabeled rather than misclassified. In other words, the hard cuts imposed were inadequate to grasp how competitive the game was and the model was more reasonable.</p>\n\n<h1>Final notes</h1>\n\n<p>In my notebooks, I have the tendency of showing almost everything I try because I write them almost for myself to learn. This time, I tried to make a selection of what is relevant and to build a story around it. In this, I partially followed the advice of @jpmiller and some professional experience. </p>\n\n<p>More importantly, I like basketball a lot and I wanted to learn a few specific things (some cross validation technique, some data wrangling, some storytelling) and I went for it. It has been a very nice experience, especially considering how much of a distraction this was from the covid situation.</p>\n\n<p>If you have specific questions, please let me know (either here or on the notebook). I will be glad to further clarify what I did.</p>\n\n<p>Cheers</p>",
  "messages": [
    {
      "id": 853039,
      "postDate": "2020-05-18T23:16:17.723Z",
      "content": "<p>This is the first time I win a competition here on Kaggle and I was definitely not expecting to end up in the top 5, so I will try my best in further explaining my solution.</p>\n\n<p>I want to start by thanking the Kaggle team for organizing the competition and providing good quality data for us to have fun with. This is by far my favorite recurring competition, despite the fact that this year's circumstances deprived us of the thrill of March Madness.</p>\n\n<p>Secondly, I would like to thank all the competitors and content creators of all 3 competitions (the two that got canceled and this one) because they have been a source of inspiration to get better.</p>\n\n<p>Here you can find my solution <a href=\"https://www.kaggle.com/lucabasa/quantify-the-madness-a-study-of-competitiveness\">https://www.kaggle.com/lucabasa/quantify-the-madness-a-study-of-competitiveness</a></p>\n\n<h1>The idea behind it</h1>\n\n<p>I took inspiration from the competition description and I spent quite some time thinking about what competitiveness would mean. The main drive for me has been exploring the play by play data, as I never used them in previous editions and I was curious to see what can I with them. I thought they were more useful than the simple game stats as it can be that stats do not tell the full story.</p>\n\n<p>I started simple and guessed that a competitive game should not finish with a large point difference, which is kind of obvious. The second step was to recognize that sometimes a team just catches fire and blows the other team away in the last part of the game. Therefore, I decided to include other criteria to account for lead changes and point differences at different stages of the game. In the notebook you will find the details and in the utility scripts you can find how they were computed.</p>\n\n<p>I then chose 6 criteria and arbitrarily decided that I was calling those games competitive. The decision was driven by the distribution of these characteristics (like the point difference at the 37th minute of the game) and get something that makes sense basketball-wise. At this point, each criteria was selecting about 10-20% of the games as competitive.</p>\n\n<h1>The model</h1>\n\n<p>The idea is to predict if the hand labeling done above is something predictable by a machine learning model, given the characteristics of the game. Given how the data is provided, I made some manipulation that can be summarized as follows:\n* Getting the total number of each statistic. For example the total number of rebounds\n* Getting the difference of each statistic. Like the difference in Steals. This is given in absolute values as I did not want the model to use who won. An alternative approach would be to duplicate the dataset (and inverting winning and losing team) as the popular raddar approach suggests (not sure @raddar is the original author of it but those notebooks are the first one I saw a couple of years ago)\n* Doing the above 2 points for different stages of the game. For example at half time, at the 37th minute, and at the end of the game.</p>\n\n<p>I wanted the model to be fairly simple as the story would have been about explaining it. At the same time, I wanted to practice a few things, namely getting partial dependence plot in a kfold validation scheme. I adopted a 5 fold validation and made sure that there was no obvious leakage of the target into the training set.</p>\n\n<p>Most of the modeling decisions were made by iterating over a Logistic Regression model. This was done because LR is simple to understand and blazing fast. At the end XGBoost showed the best performance, I don't exclude LightGBM would do better but I don't know how to tune it properly.</p>\n\n<p>The final decision was made by manually inspecting a few misclassified samples. In particular, with my great surprise, they were mostly games that were mislabeled rather than misclassified. In other words, the hard cuts imposed were inadequate to grasp how competitive the game was and the model was more reasonable.</p>\n\n<h1>Final notes</h1>\n\n<p>In my notebooks, I have the tendency of showing almost everything I try because I write them almost for myself to learn. This time, I tried to make a selection of what is relevant and to build a story around it. In this, I partially followed the advice of @jpmiller and some professional experience. </p>\n\n<p>More importantly, I like basketball a lot and I wanted to learn a few specific things (some cross validation technique, some data wrangling, some storytelling) and I went for it. It has been a very nice experience, especially considering how much of a distraction this was from the covid situation.</p>\n\n<p>If you have specific questions, please let me know (either here or on the notebook). I will be glad to further clarify what I did.</p>\n\n<p>Cheers</p>",
      "rawMarkdown": "This is the first time I win a competition here on Kaggle and I was definitely not expecting to end up in the top 5, so I will try my best in further explaining my solution.\n\nI want to start by thanking the Kaggle team for organizing the competition and providing good quality data for us to have fun with. This is by far my favorite recurring competition, despite the fact that this year's circumstances deprived us of the thrill of March Madness.\n\nSecondly, I would like to thank all the competitors and content creators of all 3 competitions (the two that got canceled and this one) because they have been a source of inspiration to get better.\n\nHere you can find my solution https://www.kaggle.com/lucabasa/quantify-the-madness-a-study-of-competitiveness\n\n# The idea behind it\n\nI took inspiration from the competition description and I spent quite some time thinking about what competitiveness would mean. The main drive for me has been exploring the play by play data, as I never used them in previous editions and I was curious to see what can I with them. I thought they were more useful than the simple game stats as it can be that stats do not tell the full story.\n\nI started simple and guessed that a competitive game should not finish with a large point difference, which is kind of obvious. The second step was to recognize that sometimes a team just catches fire and blows the other team away in the last part of the game. Therefore, I decided to include other criteria to account for lead changes and point differences at different stages of the game. In the notebook you will find the details and in the utility scripts you can find how they were computed.\n\nI then chose 6 criteria and arbitrarily decided that I was calling those games competitive. The decision was driven by the distribution of these characteristics (like the point difference at the 37th minute of the game) and get something that makes sense basketball-wise. At this point, each criteria was selecting about 10-20% of the games as competitive.\n\n# The model\n\nThe idea is to predict if the hand labeling done above is something predictable by a machine learning model, given the characteristics of the game. Given how the data is provided, I made some manipulation that can be summarized as follows:\n* Getting the total number of each statistic. For example the total number of rebounds\n* Getting the difference of each statistic. Like the difference in Steals. This is given in absolute values as I did not want the model to use who won. An alternative approach would be to duplicate the dataset (and inverting winning and losing team) as the popular raddar approach suggests (not sure @raddar is the original author of it but those notebooks are the first one I saw a couple of years ago)\n* Doing the above 2 points for different stages of the game. For example at half time, at the 37th minute, and at the end of the game.\n\nI wanted the model to be fairly simple as the story would have been about explaining it. At the same time, I wanted to practice a few things, namely getting partial dependence plot in a kfold validation scheme. I adopted a 5 fold validation and made sure that there was no obvious leakage of the target into the training set.\n\nMost of the modeling decisions were made by iterating over a Logistic Regression model. This was done because LR is simple to understand and blazing fast. At the end XGBoost showed the best performance, I don't exclude LightGBM would do better but I don't know how to tune it properly.\n\nThe final decision was made by manually inspecting a few misclassified samples. In particular, with my great surprise, they were mostly games that were mislabeled rather than misclassified. In other words, the hard cuts imposed were inadequate to grasp how competitive the game was and the model was more reasonable.\n\n# Final notes\n\nIn my notebooks, I have the tendency of showing almost everything I try because I write them almost for myself to learn. This time, I tried to make a selection of what is relevant and to build a story around it. In this, I partially followed the advice of @jpmiller and some professional experience. \n\nMore importantly, I like basketball a lot and I wanted to learn a few specific things (some cross validation technique, some data wrangling, some storytelling) and I went for it. It has been a very nice experience, especially considering how much of a distraction this was from the covid situation.\n\nIf you have specific questions, please let me know (either here or on the notebook). I will be glad to further clarify what I did.\n\nCheers",
      "votes": 6
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "853039": "This is the first time I win a competition here on Kaggle and I was definitely not expecting to end up in the top 5, so I will try my best in further explaining my solution.\n\nI want to start by thanking the Kaggle team for organizing the competition and providing good quality data for us to have fun with. This is by far my favorite recurring competition, despite the fact that this year's circumstances deprived us of the thrill of March Madness.\n\nSecondly, I would like to thank all the competitors and content creators of all 3 competitions (the two that got canceled and this one) because they have been a source of inspiration to get better.\n\nHere you can find my solution https://www.kaggle.com/lucabasa/quantify-the-madness-a-study-of-competitiveness\n\n# The idea behind it\n\nI took inspiration from the competition description and I spent quite some time thinking about what competitiveness would mean. The main drive for me has been exploring the play by play data, as I never used them in previous editions and I was curious to see what can I with them. I thought they were more useful than the simple game stats as it can be that stats do not tell the full story.\n\nI started simple and guessed that a competitive game should not finish with a large point difference, which is kind of obvious. The second step was to recognize that sometimes a team just catches fire and blows the other team away in the last part of the game. Therefore, I decided to include other criteria to account for lead changes and point differences at different stages of the game. In the notebook you will find the details and in the utility scripts you can find how they were computed.\n\nI then chose 6 criteria and arbitrarily decided that I was calling those games competitive. The decision was driven by the distribution of these characteristics (like the point difference at the 37th minute of the game) and get something that makes sense basketball-wise. At this point, each criteria was selecting about 10-20% of the games as competitive.\n\n# The model\n\nThe idea is to predict if the hand labeling done above is something predictable by a machine learning model, given the characteristics of the game. Given how the data is provided, I made some manipulation that can be summarized as follows:\n* Getting the total number of each statistic. For example the total number of rebounds\n* Getting the difference of each statistic. Like the difference in Steals. This is given in absolute values as I did not want the model to use who won. An alternative approach would be to duplicate the dataset (and inverting winning and losing team) as the popular raddar approach suggests (not sure @raddar is the original author of it but those notebooks are the first one I saw a couple of years ago)\n* Doing the above 2 points for different stages of the game. For example at half time, at the 37th minute, and at the end of the game.\n\nI wanted the model to be fairly simple as the story would have been about explaining it. At the same time, I wanted to practice a few things, namely getting partial dependence plot in a kfold validation scheme. I adopted a 5 fold validation and made sure that there was no obvious leakage of the target into the training set.\n\nMost of the modeling decisions were made by iterating over a Logistic Regression model. This was done because LR is simple to understand and blazing fast. At the end XGBoost showed the best performance, I don't exclude LightGBM would do better but I don't know how to tune it properly.\n\nThe final decision was made by manually inspecting a few misclassified samples. In particular, with my great surprise, they were mostly games that were mislabeled rather than misclassified. In other words, the hard cuts imposed were inadequate to grasp how competitive the game was and the model was more reasonable.\n\n# Final notes\n\nIn my notebooks, I have the tendency of showing almost everything I try because I write them almost for myself to learn. This time, I tried to make a selection of what is relevant and to build a story around it. In this, I partially followed the advice of @jpmiller and some professional experience. \n\nMore importantly, I like basketball a lot and I wanted to learn a few specific things (some cross validation technique, some data wrangling, some storytelling) and I went for it. It has been a very nice experience, especially considering how much of a distraction this was from the covid situation.\n\nIf you have specific questions, please let me know (either here or on the notebook). I will be glad to further clarify what I did.\n\nCheers"
  }
}