{"cells":[{"metadata":{},"cell_type":"markdown","source":"# How to Make Money Betting on March Madness using Machine Learning"},{"metadata":{},"cell_type":"markdown","source":"Even above average sports bettors lose money betting on March Madness. A bettor must pick 52.4% of games correctly just to break even, because bookies collect a portion of bets placed. Can an above average machine learning model beat the spread?\n\nI trained a machine learning model to predict the score difference between two opponents. The predictions from this model correctly picked a betting position in 65% of the games where the model identified an opportunity to bet on the spread. \n\nUsing tournament games from 2010-2013, I determined the optimal betting strategy was to bet on games where the model's prediction varied by more than 3 points from the spread. I then simulated this betting strategy on tournaments from 2014-2019. The model placed bets on 16% of tournament games from 2014-2019 where the model's prediction was 3 points more or less than the spread.\n\nThe hypothetical bets placed are summarized in the figure below. Had I bet on these 62 games, I would have made \\\\$1,580 off \\$6,820 in bets placed, a return of 23%."},{"metadata":{},"cell_type":"markdown","source":"Season | Games Bet On | Model Correct | Amount Bet | Amount Won/Lost \n--- | --- | --- | --- | --- | --- \n2014 | 13 | 77% | \\\\$1,430 | \\\\$670 \n2015 | 14 | 64% | \\\\$1,540 | \\\\$350 \n2016 | 8 | 62% | \\\\$880 | \\\\$170 \n2017 | 17 | 53% | \\\\$1,870 | \\\\$20 \n2018 | 4 | 75% | \\\\$440 | \\\\$190 \n2019 | 6 | 67% | \\\\$660 | \\\\$180 \n**Total** | **62** | **65%** | **\\\\$6,820** | **\\\\$1,580** "},{"metadata":{},"cell_type":"markdown","source":"**Introduction to Sports Betting**\n\nIn sports betting, the spread represents the difference in score between two opponents. For example, a spread of 2.5 for UVA favored against Duke means that UVA is expected to win by 2.5 points. A bettor has two options:\n\n1. **Bet favorite covers:** Bet that UVA will win by more than 2.5 points (i.e., “cover the spread”)\n2. **Bet favorite doesn't cover:** Bet that UVA will either (a) win by less than 2.5 points or (b) lose\n\nSports bettors typically must bet \\\\$110 to win \\$100. A bettor can place a bet of \\\\$110 for UVA to cover. If UVA covers, the bettor gets \\$210 back, winning \\\\$100. If UVA doesn’t cover, the bettor gets nothing back and loses the full $110. \n\nTherefore, the profit function is p x \\\\$100 + (1-p) x -$110, where p is the percentage of games picked correctly. I calculate the breakeven percentage as 52.4% by setting profit equal to \\$0 and solving for p.\n\nThe spread is set by the bookie (e.g., Vegas) and adjusts as bettors place bets in the days leading up to the start of the game. To maximize profit, the bookie adjusts the spread to entice 50% of the bettors to bet on the favorite covering and 50% to bet on the favorite not covering. If two bettors select a favorite to cover and two bettors select a favorite won't cover, the bookie collects \\\\$40 (\\$10 per \\\\$100 bet) no matter what the outcome is. Because the bookie adjusts the spread as bets are placed, the wisdom of the crowd results in the spread being an accurate predictor that incorporates all information available (team strength, injuries, etc.).\n\n**The Model**\n\nI gathered historical spread data from www.thepredictiontracker.com to utilize for this analysis. I trained a tree-based ensemble model using XGBoost that includes four features to predict the difference in score between two opponents. These four features are explained below.\n\n* Difference between Favorite and Underdog in Ken Pom Efficiency Margin prior to start of March Madness:\n\n    The Ken Pom Efficiency Margin is the offensive efficiency less the defensive efficiency, where offensive efficiency is the estimated points per 100 possessions and defensive efficiency is the points allowed per 100 possessions. Both calculations are strength of schedule adjusted to reflect efficiency against an average D-1 team (see: https://kenpom.com/blog/ratings-glossary/). If the favorite has an efficiency margin of 30 and Team 2 has an efficiency margin of 10, the difference would be 20.\n\n    I collected Ken Pom efficiency margin data from before the tournaments began using the wayback machine, which captures a historical screen shot of web pages. By relying on historical data captured prior to the tournament beginning, I prevented data leakage.\n\n\n* Difference between Favorite and Underdog in Rank prior to start of March Madness:\n\n    I took an average of all ranking systems (e.g., ESPN, AP, Massey) before the tournaments began. If the favorite has a rank of 2 and underdog has a rank of 10, the difference is 8. Kaggle provided ranking data.\n    \n\n* Difference between Favorite and Underdog in average point margin during regular season:\n\n    I took an average of the points scored less points allowed by each team during the regular season. If the favorite won by 5 points, on average, and the underdog won by 1 point, on average, the difference would be 4. Kaggle provided regular season statistics.\n    \n\n* Spread relative to favorite team:\n    \n    I used the spread as a feature in the model and transformed this feature such that the spread is always relative to the favorite (i.e., it is always positive). I also used the spread to determine who is the favorite and underdog. In the UVA example provided earlier, UVA is the favorite and the spread is 2.5.\n    \n\nAs shown by feature importance chart below, the most important feature was rank, followed by the efficiency margin. Surprisingly, the spread was the third most important feature."},{"metadata":{},"cell_type":"markdown","source":"![image.png](attachment:image.png)","attachments":{"image.png":{"image/png":"iVBORw0KGgoAAAANSUhEUgAAAZUAAADlCAYAAACbDmpOAAAgAElEQVR4Ae2d95dUVdb+/QdmXDNr1ow/iKi8OqgMCCNBRAFRsiBIUFKTQclJcmyiNE2Tc2qypE6Sc46SM0gccs40YX/Xs79z6y2Lrrea6ntv1bk+Z63qrptOePa553POPufeekUYqAAVoAJUgArYpMArNsXDaKgAFaACVIAKCKHCSkAFqAAVoAK2KUCo2CYlI6ICVIAKUAFChXWAClABKkAFbFOAULFNSkZEBSKjwMGDB+XXX3+NTOJMlQoEKECoBAjCTSoQqMCIESOkXr160qBBA/3s2LEj8JSQ24cOHZKEhISQ54VzwpEjR2Tfvn3hXBrWNUuWLJGff/45rGt5kfcVIFS8b2OWMIsKVK9eXWbPni3nz5/Xz4MHD146xrVr10qlSpVe6rqnT5++1PlunPz8+XOJj4+XLl26uJEc0zBQAULFQKMxy+4qAKisWLHid4k+efJERo0aJZ999pl+Zs6cKc+ePZO9e/dKzZo1JU+ePFK2bFlZuXKlAELvvPOOvPrqq5I7d25BTx8jn/Hjx2uce/bskTfeeEO/Y1RUv359KV26tHz33XcKsU6dOkmBAgWkXLlysnTpUgmEzeDBg7WRv337trRp00aaNm0qH3/8sZQoUUKmTp0qMTExmi5GSrh2+vTpmjeUC/GOHDlSUJ5bt27JgAEDpFChQlK0aFFZsGCBng+IoEyVK1eWhg0bann//ve/S+HChWXevHmyaNEi+fzzzyVv3rx6/OTJk1qWn376Sa9BWXAsJSVF9589e1YaN24sBQsWlJIlS6rr7tGjRxIXFydFihSR4sWLS1JSkgBgDOYpQKiYZzPm2GUF0Pj++OOP2kCjob1+/bpg5IH9ly5dktOnT0vOnDnl6NGjcvnyZQUBGsSFCxdKs2bN5Nq1a3q+/0gF4AgGFcDo6tWrCqMhQ4YIPunp6bJ582Zp0aKFoFH2D/5QAVB69+6t548ePVq+/PJLOXz4sFy4cEFhcO7cOYXK66+/LnDJAQBI78SJE7J8+XJ17128eFHLUqpUKb0WUAEkkS6ghPz4j1SOHz+umty7d0/3jxs3TgEbGxurQENZ1qxZo8C4f/++agKAALbQ5u7du+pOq1WrlpYb+c2VK5cgHwzmKUComGcz5thlBQCP7t27a6988eLFcvPmTWnfvr189NFHvnmWN998U48DMhhtNGrUSL755hv5+uuv5dSpUy8FlUGDBmmjfOPGDY2/fPny+h8jF4w+du/e/TsF/KHSqlUrHR3hhPXr12s+0NgjIC8YFWGkgjit0KRJE5k/f76OvNDYW6F58+Y6YgBUEK8Vhg4d+juoHDhwQLp166Z5xAgHIytAECMVXIsAGAEUx44dU4gBNP6hTp068sknn/j0fO2112T16tX+p/C7IQoQKoYYitmMnAKASqD7q2/fvtKjRw8dqQAk6FWjF44GHgDCiGDdunXq5sIoAN/9RyoYwQA+CBj1ZMuWTb9j37Bhw9T1A3dWhw4dJDk5WdNBGleuXNEG218Nf6i0bdtWRwU4vnHjRvn+++91RIBtpI9VYoAK3EwIGFFhhLBs2TKZNGmSYHRhuZ2wMAHuNoDBf2TiD5U7d+4oPAGl//znP1qmzp07ax6Rr+HDh2s6+PPBBx8oYOEOgz7+AQDDCAhaWnrCJcZgngKEink2Y45dViAjqKCBr1ChgjbgWNILtxj2DRw4UPr06aPuo379+glcSIAK3EyYd0CvHg3x3LlztTHGSjK4rOCOQvCHChr3OXPmSLt27XR1F1Z4AU6Y+/AP4UDlb3/7m45O4KJDOeCCQvw1atTQkQ5Wd2EOB66+QKjgWNWqVbVccP1hRAbwoSwYTbVu3TooVDBCQRmRDkZcmzZtUt3gikN6GzZsEOiJclsjLP+y8nv0K0CoRL+NmMMIKzBhwgSdWwjMxvbt26Vnz57a6KOhxMgCjSzAggnzyZMny6xZs3R0gWsxx4H9GC3APYSJc/TqAST8R8DcA9w+1mgBvfXU1FSd08G8DiAQ2NhiLgQT2w8fPlQXHOZ2EAAzpI+0EMaOHavzIhipfPvttzoiQrqYq0F6OA8LC+C+QrmseLAP8VsBUEMZcR6u3bp1q46oMHLDaAf5hbsL+fJ3YWGuB/BCmWbMmKEuRAAY7kEExIVRHlyLY8aMkcePH1tJ8r9BChAqBhmLWaUCdigAqMDlxUAFnFCAUHFCVcZJBaJYAUIlio3jgawRKh4wIotABagAFYgWBQiVaLFEFvMxceJE9afjYTQvfjA5jKfavVg2q0xeLx9siAl4q7xe/I8FGPhEe9mmTJmSxRYn+OWESnBtjDqCh++w7NWrH0z64qE6r5YP5cJEupfLhyfqExMTPV1GC5rRbke8DcGpQKg4pazL8f773/92OUV3k8PqJDy45+Xg9TcNYyk1nr73csCzOnirQrQH/4df7c4roWK3ohGKj1CJkPA2Jkuo2ChmhKIiVIQ/Jxyhumd7soSK7ZK6HiGh4rrktidIqBAqtleqSEVIqERKefvSJVTs0zJSMREqhEqk6p7t6RIqtkvqeoSEiuuS254goUKo2F6pIhUhoRIp5e1Ll1CxT8tIxUSoECqRqnu2p0uo2C6p6xESKq5LbnuChAqhYnulilSEhEqklLcvXULFPi0jFROhQqhEqu7Zni6hYrukrkdIqLguue0JEiqEiu2VKlIREiqRUt6+dAkV+7SMVEyECqESqbpne7qEiu2Suh4hoeK65LYnSKgQKrZXqkhFSKhESnn70iVU7NMyUjERKoRKpOqe7ekSKrZL6nqEhIrrktueIKFCqNheqSIVIaESKeXtS5dQsU/LSMVEqBAqkap7tqdLqNguqesREiquS257goQKoWJ7pYpUhIRKpJS3L11CxT4tIxUToUKoRKru2Z4uoWK7pK5HSKi4LrntCRIqhIrtlSpSERIqkVLevnQJFfu0jFRMhAqhEqm6Z3u6hIrtkroeIaHiuuS2J0ioECq2V6pIRUioREp5+9IlVOzTMlIxESqESqTqnu3pEiq2S+p6hISK65LbniChQqjYXqkiFSGhEinl7UuXULFPy0jFRKgQKpGqe7anS6jYLqnrERIqrktue4KECqFie6WKVISESqSUty9dQsU+LSMVE6FCqESq7tmeLqFiu6SuR0iouC657QkSKoSK7ZUqUhESKpFS3r50CRX7tIxUTIQKoRKpumd7uoSK7ZK6HiGh4rrktidIqBAqtleqSEVIqERKefvSJVTs0zJSMREqhEqk6p7t6RIqtkvqeoSEiuuS254goUKo2F6pIhVh7g/zyvZT1737+e2aLN+wzbvlO3Vdlq/f4unybT5yXlZv3+fpMq7ffVQ27DuRpTJeu/fY8WakfPnyjqXximMxM2JXFcj27geSvWOaZz9vdUqTdmOTPVs+2K7jeG+XL3+vVIlJSPG0DSsMSpFS/VOzVMZFu8/72o5Dhw5JYmKi3L9/XxYtWiTjxo2TGTNmyO3bt/WckydPyqRJk2TixIly584d33WhvhAqoRRy4fjNmzelYsWK0rBhQ+nYsaPcu3fvpVLt3LmzYGhshbt370rt2rWlQYMGUq9ePa0oT5480XhHjhyppx05ckTatm0rkydPlmPHjknjxo1l4cKFkp6ebkXj+0+omA9UQsV8G9oJFbQR/fv3FwDg6tWrsmPHDtm6dasMHjxYQYI2aNiwYTJt2jRJSEiQuLg4X3sQ6guhEkohF45fuXJFatasKc+ePRMAYujQoS+VapkyZeTEiRO+awCpunXrag8ElaNGjRoyfPhwPf78+XNNZ86cOdoDwfaAAQNk5syZgu8ZBULF/AaJUDHfhnZCZefOndKuXTtp3769QsW679HJ7Nu3r1y4cEE6deokaJvQIf3oo4+Ctg/WtdZ/QsVSIoL/YbhatWppDiZMmKCN/NOnT7Whb9SokdSpU0dSU1Pl2rVrUqxYMenatatUr15dexEYWQAqmIgdPXq0wsOCyoMHDzROVBCcf+bMGe2ZADTFixeXzz77TEaMGCF58+aVUqVKydq1a30qPHz4UDCaQbzZ/ienfNIn1bOfT2NTpefEZM+WD7brM9nb5ftqUIq0HO3tMtYbliI141OyVE8XbTsm586dUy8G7veWLVsK3FynTp1Sr8bnn38umzdvlgMHDkizZs0ULmhPChcuLOfPnxd8D/UpW7asrx2x+wvnVDKpKKBSsGBBiY+Pl6pVq+qoA6MGNOiADHoOAMvFixelQIECcvnyZdmzZ4/ExsYqaEqUKKFuM/QyAIxAqMBHGhMTIwcPHlSoYEQ0a9Ys9acii7169VKfqn92EQf8rGPHjpXX33pHvotP8eynRnyKDJ6W7NnywXbxiUmeLl/jESnSc5K3bdhhXLK0GZO1Mi5av1sGDRokn376qaDDipWdvXv3Vqj89ttv2i5g/5YtW9QlvnfvXoUOOp5wkwM+oT6lS5f2b0ps/U6oZFJOQAWjjTVr1kjlypV1VAKAwNhpaWmycuVKdY+dPn1aLIMdPXpUYQPAYNSBkQhGM4BRIFQCRyqZgYp/1un+Mt91QveX+Ta0y/2FdmTTpk2SkpKi7crGjRvl0qVLestjX/PmzeXGjRva2QRUcByd2swGur8yq5SD5/m7v7AiI0+ePHL48GF1cwEIGDEAOughWENLCyqoDDi2ZMkSadq0qY5y/KGClR1wrWHS7fr165keqfgXl1Axv0EiVMy3oV1Qse5ttC0dOnTQeZNChQpJrly5tC0BdNA5Xbp0qeTPn19HM2fPnrUuC/mfUAkpkfMnAAL9+vXzJdSzZ0/BRBqW+zVp0kTnWLAqA6MXTJ4hwC86ffp07VFgH1Z/YfUGJt2x/A8T/1j5hQl7xIPJNrjBsBAAI5VVq1bJ8uXLNa6pU6fKhg0bfOkHfuHDj4GKmLfNhx/Ns1lgjvnwIx9+DKwTxm4TKsaazpdxQsUnhbFfCBVCxdjKG5hxQiVQEfO2CRXzbBaYY0KFUAmsE8ZuEyrGms6XcULFJ4WxXwgVQsXYyhuYcUIlUBHztgkV82wWmGNChVAJrBPGbhMqxprOl3FCxSeFsV8IFULF2MobmHFCJVAR87YJFfNsFphjQoVQCawTxm4TKsaazpdxQsUnhbFfCBVCxdjKG5hxQiVQEfO2CRXzbBaYY0KFUAmsE8ZuEyrGms6XcULFJ4WxXwgVQsXYyhuYcUIlUBHztgkV82wWmGNChVAJrBPGbhMqxprOl3FCxSeFsV8IFULF2MobmHFCJVAR87YJFfNsFphjQoVQCawTxm4TKsaazpdxQsUnhbFfCBVCxdjKG5hxQiVQEfO2CRXzbBaYY0KFUAmsE8ZuEyrGms6XcULFJ4WxXwgVQsXYyhuYcUIlUBHztgkV82wWmGNChVAJrBPGbhMqxprOl3FCxSeFsV8IFULF2MobmHFCJVAR87YJFfNsFphjQoVQCawTxm4TKsaazpdxQsUnhbFfCBVCxdjKG5hxQiVQEfO2CRXzbBaYY0KFUAmsE8ZuEyrGms6XcULFJ4WxXwgVQsXYyhuYcUIlUBHztgkV82wWmGNChVAJrBPGbhMqxprOl3FCxSeFsV8IFULF2MobmHFCJVAR87YJFfNsFphjQoVQCawTxm4TKsaazpdxQsUnhbFfCBVCxdjKG5hxQiVQEfO2CRXzbBaYY0KFUAmsE8Zu582XT24/SPfs59b9x7J916+eLR9st3X7Tk+X7+LVm3Lg8FHPlPHuoye+9uLevXty584dOX78uFy8eFGePHmi23fv3tXvOPHRo0e6D+fheCRD+fLlHUv+Fcdi9mjE169fl/3798vTp0/l1KlTgkoTTli6dKmsXLlSP3v37pVnz55pnCdPnpT79+9rpdu8ebOg93rjxg1ZtWqV7N69W9LT0zNMLtu7H0j2jmme/bzVKU3ajU32bPlgu47jvV2+/L1SJSYhxTM2LDFkrd6LAESPHj2kVatWEhMTI4sXL5YRI0ZIkyZNpHnz5pKWlqb39owZM3S7cePGMm7cuAzvY7d2EipuKZ2JdLZs2SLt27fXRh9wQYMfTsiZM6dMnDhRpkyZIm3atJHZs2drnHv27JHbt28LhtE1atRQ6CQlJUndunUVLOjtZBQIFfOBSqiYZUMLKrgf0dl7/vy5JCYmSkJCgnz44Yd6Hx87dkx69eql3y9fvqz3+Pnz56VQoUIZ3cau7SNUXJP6fxNq0aKFVK1aVfr3768VAo3+N998I99//732Nh48eKC9jcOHD+twF70SQGDu3Lny+PFj6datm3Tq1EmqVKmiPZUuXbpI7dq1BVBC+Ne//qXxoiJidPL+++8LRkHDhg3TEVDZsmUlR44cguuKFi0qefLk0d7P/+bw998IFbMapIxGlYSKWTb0hwo8DRUrVpRixYrJihUrBHOccIXt2LFD0IBfvXrVd8OuWbNG2wLfjgh8IVQiIPqFCxe0kQcY1q5dq5Xk1q1bEh8fLw0bNhT4UGNjY9U9hSEvXFinT5/WUcxvv/2mw2AAAqOZXLlyyZEjR3TUERcXJwCSBRWraB9//LG6tzp37qzn4rqWLVvq4TFjxsjw4cOtU33/z549K82aNZPixYvL69myy0/Tkjz9mTLX2+Wb9rO3y5cwPUnGz/ZOGUfMSpUNGzb4PuvXr5fevXtrZxJeiE8++UTKlCkjRYoU0Y4lzkWns1y5cuqZ8L/W7e9ffPGFrx2x+wvnVDJQFC4tjEhat26tDTYqChpvBPQy4K6yoIKRR/78+aVt27b6wTUAUqNGjbSyXbt2TUqVKqXXotcCqABO/lDB/AxGIgDPy0DFP+scqZjVy+VIxXx7WSMVzKlYc52YT+nZs6fcvHlTb899+/ap+wuT85jAx9xLamqqzrH4379ufzd2pAKxMeyL9EqHlzUYehxwWwEcgAUm3dDjePjwoUybNk2BY0EFoKhQoYKg8iCgMqGCASqYaIdLCz0TBJw7ePDg30EF2syfP1/jB2wIlYwbG07UZ6xLRnCK1n1enai/dOmSDBo0SPr166edT9zP8+bNkz59+uioBe4wdBxxHKMXgAXnRzIYBxVMZGNVFCaY0Rijp29SABiQb8yRAAgzZ87UCgKfaf369eWHH37wjVSwImvnzp3y7bffSqVKlaRdu3aC0UkoqLz99tsKo6+++kowf4OVZFj1Rahk3HgSKhnrEq0AyShfXoUKOpFweWMOZdu2bXLu3DnBpDwm6c+cOaNLidH+wV0NbwTOO3HiRESbROOggklma/4Bq5nQQHstoNJgMv7o0aNRUTQ+/BgVZshSJvjwY5bki4qL+fCjQw8/Ym325MmTpW/fvnLlyhWpXLlyVBjczkxgzTlcYZYv1c64w4mLUAlHtei6hlCJLnuEkxtCxSGoYBUUfIpYbouVTpjcZnBWAULFWX3diJ1QcUNlZ9MgVByCCnyKWGY7dOhQXfEwdepUZy3J2HXJs5dlwPM8cKV6ORAq5luXUHEIKtWqVdMHgDDngGAtqTW/ykRvCThSiV7bZDZnhEpmlYre8wgVh6CC5bhY4QCoYEktltwyOKsAoeKsvm7ETqi4obKzaRAqDkFlwYIF0qFDB32SFE+FT5gwwVlLMna6vzxQBwgV841IqDgAFfi+8bAPJuvx6hKMWPDQIIOzCnCk4qy+bsROqLihsrNpECoOQAUmq1Wrlj7Ih5esWR9nTcnYCRXz6wChYr4NCRWHoIIn6vGsClZ94Wl0fBicVYBQcVZfN2InVNxQ2dk0CBWHoJKSkqJv4cTr4q2Ps6Zk7ISK+XWAUDHfhoSKQ1BZtmyZLFq06Hcf86tLdJeAUIlu+2Qmd4RKZlSK7nMIFYegMmTIEF1OjJcj4iWM+DA4qwCh4qy+bsROqLihsrNpECoOQcXfbPj5W7zBl8FZBQgVZ/V1I3ZCxQ2VnU2DUHEIKvjddfxSIX75EL/PjNVgDM4qQKg4q68bsRMqbqjsbBqEikNQWb58uf6WSnJysv4UL37EhsFZBQgVZ/V1I3ZCxQ2VnU2DUHEIKoE/yhW47axZ/5ixEyrm251QMd+GhIrNULl9+7aOTPCb7Ril4LNw4ULBrxsyOKsAoeKsvm7ETqi4obKzaRAqNkMFL48ERPBW4okTJ+oHP9a1cuVKZy3J2PnuLw/UAULFfCMSKjZDBe/9evr0qf5G++PHjwUC43eZz58/b35tifIScKQS5QbKRPYIlUyIFOWnECo2Q8WyN0YrBQsWlOzZs8u7774ruXPntg7xv0MKECoOCetitISKi2I7lBSh4hBUypUrp28n7tq1q1y4cEEaNmzokAkZraUAoWIpYe5/QsVc21k5J1Qcggp++REuL7xUEnMqJUuWtDTnf4cUIFQcEtbFaAkVF8V2KClCxSGo4Dfq09PTZefOnfoA5K5duxwyIaO1FCBULCXM/U+omGs7K+eEikNQuXv3rmzcuFG2bNmiP9B18uRJS3P+d0gBQsUhYV2MllBxUWyHkiJUHIJKbGys/pxwx44dBc+uVKpUySETMlpLAULFUsLc/4SKubazck6oOASVChUq6JxKt27d9JcfS5cubWnO/w4pQKg4JKyL0RIqLortUFKEikNQwWtZ+vXrJ3Xr1pUBAwZI69atHTIho7UUyP7PXFKw/0rPfj4esFK6T07zbPlgu95TvF2+snHLpfnYxY7Y8PaDdOtWiOh/QsVmqFi/R4+HIKdNm6arv/CmYsyxeCVcuXJFTp8+LWfOnJHLly/rw55Xr1717btz544jRYW2eDEntM0oZHv3A8neMc2zn7c6pUm7scmeLR9s13G8t8uXv1eqxCSkOGLDm/fT5dixY/rbTWXKlNGOLO7TRo0aSfny5QUrUnH/4J7t1KmT7mvSpIlgUZGdgVCxGSow6tGjR9VGGzZssNNWURNXy5YtFZYYjc2YMUOOHDmi80c9e/aU7t27y7p169Tl939lOBgY/q9rMDcVExMjDx48yPA0QsV8oBIq4dsQUEHH68mTJ7rytESJErJnzx79jhsGPxiIV0clJSXJiBEjdD9eHxUfH5/h/RTuTkLFZqjgt+nxEkmEKlWqhGuXqL6uTZs2cujQIV8e8bPJ+M0YBCyjxo+S4R1oP/74o8TFxakbEA+A3rt3T/r06aMVeunSpfr6Gixo6NGjh+zfv19HPGlpadKlSxfBfkAEN8jixYsFD5GOHDlS6tSpQ6h4eDRGqGQNKrgHseq0cePGeq/cvHlT7y0sGPr666/l8OHDgjYK9yVeIzVnzhxp2rSp3me+GzqLXwgVm6GSmprqg0rZsmWzaJ7ovBwPdBYqVEgf6Bw9erRgufSnn36qlRYjF8AAD37+9a9/le3bt8uECRO0Et+4cUOKFy+uI5n79+9LjRo1FE67d++WBg0aaC8LQ3O40qZPny79+/eXgwcPKmQAscTERMGbCvxHKhi6Y74KL/B8PVt2GTA1ydOfyXO9Xb4pHi/fkMQkGTfbGRuu27BRNm3apC+vnT17thQtWlQWLFgga9asEXT80MnFqGTJkiX6ho/ChQsLFhThV2nhXcC1dnxWr16tadoRl5NxfPnll441sK/YGTOMh/d85c2bV95//33Jly+ffsd/r4TAkQrK9fDhQ/nll1/UPQW32IkTJyRPnjxaZLgD69evL4AKRhqYXwJUXn31VfX/VqxYUQBgjEoAIPSoihUrJpUrV5ZVq1bpw6PoVd26dUuqVq36O6j4a0r3V/i93GiZi+JIJXwbwv2Fl9nCtYwPvAfonGEfAn44sFWrVr7j2AeYwCtgZ+BIxeaRip3Gida4AqGCkQkggYCHPS13Vo4cOQTDbwzHcQ2gUq9ePYUCXGQArfWLmDgPE/yAD97qPG/ePJ1IRE9l4MCBeuzAgQMKHP+Rir9GhEr4DRKh4o52Tk/UY2Q/c+ZMmTVrlnbgcD9OnTpV5s6dK5gLRcfv2rVrAo8KzoFbGfMudgZChVB56foUCBW4rzAP0r59e2nRooXMnz9fTp06JYAK5lCw+gQTgv5QQaJYHYfzcd2oUaN0tAPfL9xZ+GDFClxhWJLdtm1bwQioevXqQUcqH+T+UFL2/MfDn/OStGqzh8v3H0laudHT5UvZcVKS1+9ypIzpT5/pT23AvYV5yL1796pXYNmyZbqNDhq8AfAUYBER5i/37dsXdDXlSzcM/72AUCFUXrruYESBymkFTM4DGOgB4T+2MacCny1cVthnDcsxqrFWfiEO6zprGTIm8xEPtjF6wbkYBWHiH/v8r7fSt/7z4UdLCXP/8+FHc21n5ZxQIVSsumDrf0y416xZ09Y4Q0VGqIRSKPqPEyrRb6NQOSRUCJVQdcSY44SKMaYKmlFCJag0xhwgVAgVYyprqIwSKqEUiv7jhEr02yhUDgkVQiVUHTHmOKFijKmCZpRQCSqNMQcIFULFmMoaKqOESiiFov84oRL9NgqVQ0KFUAlVR4w5TqgYY6qgGSVUgkpjzAFChVAxprKGyiihEkqh6D9OqES/jULlkFAhVELVEWOOEyrGmCpoRgmVoNIYc4BQIVSMqayhMkqohFIo+o8TKtFvo1A5JFQIlVB1xJjjhIoxpgqaUUIlqDTGHCBUCBVjKmuojBIqoRSK/uOESvTbKFQOCRVCJVQdMeY4oWKMqYJmlFAJKo0xBwgVQsWYyhoqo4RKKIWi/zihEv02CpVDQoVQCVVHjDlOqBhjqqAZJVSCSmPMAUKFUDGmsobKKKESSqHoP06oRL+NQuWQUCFUQtURY44TKsaYKmhGCZWg0hhzgFAhVIyprKEySqiEUij6jxMq0W+jUDkkVAiVUHXEmOOEijGmCppRQiWoNMYcIFQIFWMqa6iMEiqhFIr+44RK9NsoVA4JFUIlVB0x5jihYoypgmaUUAkqjTEHCBVCxZjKGiqjhEoohaL/OKES/TYKlUNChVAJVUeMOU6oGGOqoBklVIJKY8wBQoVQMaayhsoooWwZKhIAABJjSURBVBJKoeg/TqhEv41C5ZBQIVRC1RFjjhMqxpgqaEYJlaDSGHOAUCFUjKmsoTJKqIRSKPqPEyrRb6NQOSRUCJVQdcSY44SKMaYKmlFCJag0xhwgVAgVYyprqIzmeD+31Jm0zbOfmMnbZMic5Z4tH2yXMGdZlss3d8dZuXXrlnTu3FkaNWokAwcOlMuXL8uMGTOkadOm0qZNG9m5c6c8f/5cLl26JD169JDGjRvL4cOHQ1WxLB+/c+eOHD9+PMvxRHMEhEoUQGX37t3yxhtvyIkTJ16qrrz77ruSL18+KV68uOzZsyfDa9etWycLFy6UZ8+evXD8wYMHsmzZMrl//74eQzwtW7b0nVehQgWpV6+eb/tlv0yYMEGQfkYBN/U//vEPweji448/lqlTp+ppp06dkilTpsiNGzdk165dkj9/fhk5cqQkJiZK3rx5ZdiwYRlFp/uyvfuBZO+Y5tnPW53SpN3YZM+WD7brOD7r5Ru45LBcv35dtm/fLqjjCQkJMnPmTN33+PFj2bp1q3To0EHS09O1ro0fP1727dsnRYsWDVq37DpAqNilZNbjKV++fNYjCRLDK0H2u7IbvaUhQ4ZIrVq1pGvXrnL79m25cOGCQuDJkydy5swZvTF+++03BcehQ4d8xz/66CNBJV2yZIn88MMPeg3OB6ROnjyp+cfNhd7Yo0eP9Hr0xnADPXz4UHtMNWrUkDVr1mgauKnQY0NaZ8+elRIlSkjt2rU1XvSuAK6jR49qXLg5kc6RI0c0P1evXhW4LpA/pI3jiAPpozyIc+/evXo9juGG//bbbzWPV65c0e8///yz5uP8+fNy7do1adGihfz00096DfICACK+YIFQMR+odkHFqiPoTE2fPl2mTZum9RgdN3RQevXqJbi/WrduLfv379fT0YFBXXUyECpOqvtycXsWKmgk27dvrxX7iy++kPXr18uAAQPk5s2bcvHiRWnXrp32rJo3b669dAzfcRxQsKCyevVqhQEaYwztBw8eLA0bNtTGPTk5WcaNG6dwevvtt7XXjziWL18uGzZs0N5Z//79BQ176dKlNe5ffvlFG3PcjIAKenTo6Q0fPlxdBcgj0nrzzTdl1KhROhrp2LGjIJ7evXtL3bp19ebEjYv0Mbpo1qyZXo/y7Nix43dQQVVAnGXLllUoIR5Ap0yZMuqumD17tuTJk0fLBWj5B7g5UlNTZfLkyfL6W+9ItbgUz36qD0mRn6Yle7Z8sN2QxKQsl29M2lY5ffq0frZt26Z1LyUlRbcxavnuu+8kLi5OOzro2GzcuFGPFStWTFauXOm71orDzv/olKH+2xlntMWFziNAHW35CswP2junQkRHKps3b5ZBgwbpiAMjlXnz5unQHL39xYsXa+M+adIkmThxovqA0ZvH+YAK3F/169eXBg0aCMCCxhXHnj59KqNHj5b4+Hht1C2ooGHGdUlJSXoegNaqVSsFCsQFueF3hrvpm2++0ZEGoIL40NODf7pkyZLqTsCI6L333lPXGW4UgATwwSgIYEOPzx8qcDEgHvQSASj/kQrSxgiqQIECPqgAcrgeeUUAYDB6CQwoz4EDBxS82XLklM9iUz37KdY3VXpNSvZs+WC72MlZL1/CL7+q+xR1qEmTJlqnMacClyo+GF3HxMRop+37779XqGA/3Kuoh9Z5TvxHZwz3iBNxR0ucaA/QfkVLfoLlA51Yp0LEoIKhORrb1157TXLkyKHzKmjksW/WrFk62sC8Ahp5zIsgADSABRpTzEdAMAzj4UZDjx4wQQCcevbs+TuoWD5jjFIwosgIKlu2bJGaNWtKt27dFDaACnzQABf80+jxDR06VDD3UaRIEU0LbjGMpBDgJgMcA6GC65DHuXPn6vxJIFQWLVqkLkCMRKyRU2agoon+9w/dX3R/YV4Gcyqoa+gEwbWM7wjWnCXuKRyDGxadNXTaAJrChQv7VydHvtP95YisYUXqSfcX5k/Q8GM4jICeeNu2bWXFihVSpUoVndNA7x4NPRptTG53795dYmNjdV7Dcn9Zih47dkyH+hj9YAQDGPi7vwKhAtcRJiytyXqIjBsPDT56VOjpASrYhvsK6cNVh0bfHyp3794VzM2sXbtW/dd16tTJFFRKlSql8zAob7Vq1XSOhlAJDgZO1AfXxn+BBqCCOTy4szAiwSganRncO6jPGJ3ALQbYoPODOo2OFEYQTgdCxWmFMx+/J6Fy7949dVkBLgjoOWHSHA02RhyYgEfASAQ9eYxYMCeBnhVgM2bMGIWLniSi7ic00CNGjNB4sR/DeUALAMF1CAAH3GWId9WqVXo+AII5FP+JcORv/vz5mq85c+bo3AxcV4Ac5nywussKmLSHvxrAwUgF8SANpI/z0RPETQxfKyb0Aa1+/fpp2igreo8IyAdWjCFtXG/NocBtZq1Ss9IM/J8zVx6JW3bUs58hy47ItNS1ni0fbJeYuibL5Vt/7Gpg1YiabUIlakyh7n6nchMx91dmCwTYoKHG5DtWRLnRo8ps3qzz4HqDSwG9QoyOACy3Ax9+dFtx+9Pjw4/2a+p2jHxOJQqeU8mM0dHLxxyM5R/OzDVunmPlL6PnYdzKB6HiltLOpUOoOKetWzETKoZAxa0KYXI6hIrJ1vv/eSdUzLchoUKomF+L/1sCQsV8UxIq5tuQUCFUzK/FhIpnbEiomG9KQoVQMb8WEyqesSGhYr4pCRVCxfxaTKh4xoaEivmmJFQIFfNrMaHiGRsSKuabklAhVMyvxYSKZ2xIqJhvSkKFUDG/FhMqnrEhoWK+KQkVQsX8WkyoeMaGhIr5piRUCBXzazGh4hkbEirmm5JQIVTMr8WEimdsSKiYb0pChVAxvxYTKp6xIaFivikJFULF/FpMqHjGhoSK+aYkVAgV82sxoeIZGxIq5puSUCFUzK/FhIpnbEiomG9KQoVQMb8WEyqesSGhYr4pCRVCxfxaTKh4xoaEivmmJFQIFfNrMaHiGRsSKuabklAhVMyvxYSKZ2xIqJhvSkKFUDG/FhMqnrEhoWK+KQkVQsX8WkyoeMaGhIr5piRUCBXzazGh4hkbEirmm5JQIVTMr8WEimdsSKiYb0pChVAxvxYTKp6xIaFivikJFULF/FpMqHjGhoSK+aYkVAgV82sxoeIZGxIq5puSUCFUzK/FhIpnbEiomG9KQoVQMb8WEyqesSGhYr4pCRVCxfxa/N8S5MyZU9atW+fZz9q1a2XixImeLR9sN378eE+Xb8mSJTJz5kxPl3H+/PmyaNGiqC9jkSJFHGv7XnEsZkbsqgLvvfee3rC4ab34mT59upQrV86TZbPsVbJkSU+Xb8iQIVKvXj1Pl7Fdu3bStWvXqC/juHHjHGufCBXHpHU34qJFi7qboMupPXnyRH744QeXU3U3ufr167uboMupHT58WOLj411O1d3kMFJZtmyZu4lGWWqESpQZJNzsfPnll+FeasR1gEqrVq2MyGu4mWzSpEm4lxpx3ZEjR2TYsGFG5DXcTML1tWLFinAv98R1hIonzCiyatUqj5Qk42I8f/5cdu7cmfFBj+zdtm2bR0qScTFu374tAIuXw7lz5+TixYteLmLIshEqISXiCVSAClABKpBZBQiVzCrF86gAFaACVCCkAoRKSImi+4S9e/dK2bJlpWDBgjJ79myBm8jUcOjQISlRooRgJRvCo0ePZOTIkVKoUCEpX768XLt2TZ4+fSqDBw+WfPnySe3ateXs2bNGFRcuoLZt26q9sJoNS6WvXr2qK9tQTpQX5X78+LGWL3/+/LqU2iS7pqenq72wbLVq1aqyf/9+gW2xDfuuX79e6+mBAwekVKlSUrhwYSNdmwcPHpRPP/1Ulw/DrYfyff755+qKhr2OHTsmpUuX1vJt2bLFqHqalcwSKllRL8LXouHp3bu3bNy4UdBYtWjRQs6cORPhXIWf/L179wQNTaVKlTSSffv2SZ8+feTKlSsyY8YMGThwoOzYsUOwbBM3LVbajBkzRp49exZ+oi5f+fDhQwUh8r9161bp3r27dO7cWZegopwoLxrhSZMmyaBBgxSkMTExcufOHZdzmrXkUE6Ucdq0aTJhwgSFBx7uxLxRz549BcdRdkxqo3FG42tSQP47dOggjRs3VqhUrlxZwYgyonx3797V8uHZHKx6++qrr0wqXpbySqhkSb7IXnz58mWFyqlTp7RnGxcXp73AyOYqa6nfuHFDqlSpopGgF49RCaBx69YtqV69uo7G0FAhbN++XRve+/fvZy3RCF2NBx5jY2O1V49OAcqJ8mJ/jRo1xOrdorMA0JgUUBYAsmHDhpKWlib//Oc/FTLXr1+Xpk2baqOLUQzqMOCD0alJdsSDqnjWY+jQobqE+MMPP9RyoP7CXhhVV6xYUSftUT6MrE0qX1bqGqGSFfUifC1eCYEbF6MTuBywXNP0VWD+UEFZAEoElO/rr7/Wni/cfAi7du1SqKBXaFpAA1O3bl0FR4UKFQRLphFQ3jVr1gj2wbWJ8OOPP2oPXzcM+oMeOhrdxMREyZUrl+YcIy6ABjZDGW/evKn78+TJ4/se7UXEyArgACBRvtTUVClQoIBmG+XC81RwaeJhVtRnBLg2rbJGe/mymj9CJasKRvB6VFIMtY8eParuBPR6TV926w+VTZs2KTTgajh+/Lg0atRIb+ARI0ao6vDNo2cP4JgU0LBijgjPNCCgkUX5UE64vDZv3qzP5Pzyyy86esFT6OfPnzemiOiZ44MA1yzcltmzZ5cHDx6o669169baa8fDnla533nnHd810V7QBQsWCOa6AMLXX39dypQpI2+++aaWDx299u3bK0zq1KmjS6gxR5Y7d25jypdV/QmVrCoYwetx48L3Pnz4cJ1fgG/eNN+7v3xwcSUnJ8snn3yiLhMAsl+/fjJr1ixp06aNwD+NmxYPQcKlgsYKDa9JASOU5s2by3fffSfLly+X3bt3S1JSkpYP5ezfv7+WEb559HjxChfMIVkjGRPKil56SkqKLF26VG2EkSXKhQ5AQkKCLjxAeVA27McIGyNu0wJsiZEK3JV4BQ06BOjwwC0GkMyZM0fdm1h8YWL5wrUHoRKuclFyHXr2aJTQIJ04cSJKchVeNuCHxuQ7bkq4TDBXhLkENEpopKwRCRpiTNzjdRim+anhHpk3b542OHPnzvUtskD5UE6U11p4AOig4YUbyer5h6esu1ehY4PePBpVdAQwosa+n3/+WfdbDwdiP0Zr0AGuJNMCwAgXJTo6VvlQf7GNgHky3JsoHxZh/FECofJHsTTLSQWoABVwQQFCxQWRmQQVoAJU4I+iAKHyR7E0y0kFqAAVcEEBQsUFkZkEFaACVOCPogCh8kexNMtptAJYkoqH6fDBU+gmTdwbLTwz/9IKECovLRkvoALuK/CXv/wlw0QBF6wywjJerATEu9GwSg7b+GB1HM7BMWxjlRK2sQoNq6+w4g6v+8F+HMc+xMFABcJVgFAJVzleRwVcVODPf/6zLF68WEcpAIQV8FAk3kHVsWNHfU4C21iO3bJlS32nGEY1WPaKByzxvAvet4WXcOJBWbyQE8+O4CFT7O/SpYs+uIcl2xwJWQrz/8sqQKi8rGI8nwpEQIE//elP+owHXgmC0YUV8HZcvGkAz3tcuHBB8CNR+AVJnIPnXfAsBX4zHQ+LYkQCiOD5F0AFD+ThnLFjx+oT/r169ZJq1arpc0LWszJWOvxPBTKrAKGSWaV4HhWIoALB3F+ABt5Fhbc116xZUx+2w4jEelAUWcaT+XiAEmHUqFEKGEBl+vTpuq9v374ydepUfYfc6dOnfe+r0oP8QwVeUgFC5SUF4+lUIBIKBIMKnkTHC0UvXbokgAOe4G7QoIG+PwxPdmPEAmB069ZN37gANxdePw+o4K0ECKtXr9Y3QMN1hvfIWU+8R6KcTNN8BQgV823IEvwBFMDvdmQUMD+CNxtjPgSvt8Fv0pw8eVLnSAASvPEYL6rE+7VwDl6fgm3AB3MpCJg/wetFOnXqpNdh5MNABcJVgFAJVzleRwWoABWgAi8oQKi8IAl3UAEqQAWoQLgKECrhKsfrqAAVoAJU4AUFCJUXJOEOKkAFqAAVCFcBQiVc5XgdFaACVIAKvKAAofKCJNxBBagAFaAC4SpAqISrHK+jAlSAClCBFxQgVF6QhDuoABWgAlQgXAUIlXCV43VUgApQASrwggKEyguScAcVoAJUgAqEqwChEq5yvI4KUAEqQAVeUIBQeUES7qACVIAKUIFwFSBUwlWO11EBKkAFqMALCvw/dVx/hMaiiS0AAAAASUVORK5CYII="}}},{"metadata":{},"cell_type":"markdown","source":"I compared this model against prior Kaggle March Madness competitions to evaluate the predictive performance. Using the four features above, I instead predicted the probability of the favorite winning. This allowed me to compare the model’s performance to the historical Kaggle leaderboard, which ranked entries based on log loss of predicted win probabilities. This model would have performed in the top 19% of all March Madness models from 2016-2019. It would have placed in the top 1% once. This confirms the model is above average.\n\n**Betting Strategy**\n\nAfter the model generated predictions, I determined the optimal betting strategy. The model is more confident if the predicted score difference is further from the spread. For example, the model could place bets on games where the prediction differs by more than 5 points compared to the spread. In the UVA example above, the model would bet when the predicted score difference is greater than 7.5 (2.5 + 5) or less than -3.5 (2.5 – 5). Although a higher threshold suggests greater confidence, the model may place bets on fewer games, resulting in lower returns.\n\nIn the 10 games from the table below, two games differ by more than 3 points. If the threshold was set to 3, the model would have bet on these two games only. "},{"metadata":{},"cell_type":"markdown","source":"| Season | Favorite | Score | Underdog | Score | Actual score Diff | FSpread | Predicted Score Difference | Predicted - Spread\n--- | --- | --- | --- | --- | --- | --- | --- | --- | --- | ---\n681 | 2013 | north carolina st. | 72 | temple | 76 | -4 | 4.0 | 4.12 | 0.12\n687 | 2013 | saint louis | 57 | oregon | 74 | -17 | 4.0 | 2.19 | 1.81\n689 | 2013 | gonzaga | 70 | wichita st. | 76 | -6 | 6.5 | 7.29 | 0.79\n**691** | **2013** | **san diego st.** | **71** | **florida gulf coast** | **81** | **-10** | **8.0** | **11.23** | **3.23**\n695 | 2013 | mississippi | 74 | la salle | 76 | -2 | 4.0 | 2.23 | \t1.77\n**698** | **2013** | **miami-florida** | **61** | **marquette** | **71** | **-10** | **5.5** | **2.26** | **3.24**\n700 | 2013 | indiana | 50 | syracuse | 61 | -11 | 5.0 | 5.73 | 0.73\n705 | 2013 | kansas | 85 | michigan | 87 | -2 | 2.0 | 1.07 | 0.93\n707 | 2013 | ohio st. | 66 | wichita st. | 70 | -4 | 4.5 | 5.73 | 1.23\n709 | 2013 | florida | 59 | michigan | 79 | -20 | 3.0 | 5.79 | 2.80"},{"metadata":{},"cell_type":"markdown","source":"I determined the optimal betting strategy using games during 2010-2013 and applied this strategy to \"unseen\" games from 2014-2019. I then evaluated the hypothetical winnings had various strategies been applied. I tested different thresholds of the difference between the spread and model predictions in games from 2010-2013 for six different thresholds, ranging from 0 to 5. The winnings from 2010-2013 for each threshold are summarized in the table below."},{"metadata":{},"cell_type":"markdown","source":"Model Difference from Spread | Amount Won/Lost | Games Bet On\n--- | --- | --- \n0 | -\\\\$3,330 | 246\n1 | -\\\\$2,430 | 171\n2 | -\\\\$1,410 | 114\n**3** | **\\\\$690** | **72**\n4 | -\\\\$20 | 46\n5 | -\\\\$260 | 31"},{"metadata":{},"cell_type":"markdown","source":"The only threshold with positive returns was the optimal threshold of 3 points, which returns \\\\$690 from 2010-2013. Therefore, I applied this betting strategy to an \"unseen\" set of games from 2014-2019. When the model returned a predicted point difference that was more than 3 points different than the spread, the game was selected to bet upon and the model picked whether the favorite would cover. Consider the examples below.\n\n(721) Louisville was the favorite against Manhattan and won 71-64, a difference of 7 points. Louisville was favored to win by 17 points, as represented by the spread. The model predicted Louisville would win by 11.7. Therefore, the model predicted Louisville would not cover. Because Louisville did not cover the spread, the model won \\\\$100 off \\$110 bet (i.e., bet \\\\$110, returned \\$210).\n\n(732) Wisconsin was the favorite against American and won 75-35, a difference of 40 points. Wisconsin was favored to win by 14.5 points, as represented by the spread. The model predicted Wisconsin would win by 11.4 points. Therefore, the model predicted Wisconsin would not cover. Wisconsin won by far more than 14.5. Therefore, the model lost the full \\\\$110 bet (i.e., bet \\$110, returned \\\\$0)."},{"metadata":{},"cell_type":"markdown","source":"| Season | Favorite | Score | Underdog | Score | Actual Score Difference | Favorite Covered | FSpread | Predicted Favorite Covered | Predicted Score Difference | Amount Bet | Amount Won/Lost\n--- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | ---\n**721**\t| **2014**\t| **louisville** | **71** | **manhattan** | **64** | **7** | **0** | **17.0** | **0** | **11.75** | **\\\\$110** | **\\\\$100**\n723\t| 2014 | michigan st. | 93 | delaware | 78 | 15 | 0 | 15.0\t | 0 | 11.69 | \\\\$110 | \\\\$100\n**732** | **2014** | **wisconsin** | **75** | **american** | **35** | **40** | **1** | **14.5** | **0** | **11.45** | **\\\\$110** | **-\\\\$110**\n733 | 2014 | arizona | 68 | weber st. | 59 | 9 | 0 | 20.5 | 0 | 15.44 | \\\\$110 | \\\\$100\n740 | 2014 | memphis | 71 | george washington | 66 | 5 | 1 | 4.0 | 0 | 0.63 | \\\\$110 | -\\\\$110\n748 | 2014 | wichita st. | 64 | cal poly slo | 37 | 27 | 1 | 17.0 | 1 | 22.20 | \\\\$110 | \\\\$100\n754 | 2014 | michigan st. | 80 | harvard | 73 | 7 | 0 | 8.5 | 0 | 3.52 | \\\\$110 | \\\\$100\n772 | 2014 | michigan st. | 61 | virginia | 59 | 2 | 0 | 2.5 | 0 | -2.47 | \\\\$110 | \\\\$100\n776 | 2014 | kentucky | 75 | michigan | 72 | 3 | 1 | 2.0 | 0 | -2.47 | \\\\$110 | -\\\\$110\n778 | 2014 | kentucky | 74 | wisconsin | 73 | 1 | 0 | 1.5 | 0 | -2.24 | \\\\$110 | \\\\$100"},{"metadata":{},"cell_type":"markdown","source":"**Evaluating Betting Strategy**\n\nWhen I applied this strategy to games from 2014-2019, the model identified 62 games where the predicted score was more than 3 points different than the spread. I simulated placing \\\\$110 bets to win \\$100 if the model picked correctly. If the model picked incorrectly, the full \\\\$110 was lost. The model placed the proper bet in 65% of games, winning \\$1,580 on \\\\$6,820 bet. The winnings were consistently positive in all tournament years where the strategy was tested.\n\nThe model more often selected the favorite not to cover, but it was more often correct when it selected the favorite to cover. During the simulation applied to 2014-2019, the model predicted the favorite would not cover in 82% of games and would cover in 18% of games. The model was correct in 59% of the 51 games where it picked the favorite would not cover. It was correct in 91% of the 11 games where it picked the favorite would cover. The suggests the model more often identified games where the favorite is overvalued (the spread is greater than the prediction), but is more accurate in games where the model determined the favorite is undervalued (the spread is less than the prediction)."},{"metadata":{},"cell_type":"markdown","source":"Season | Games Bet On | Model Correct | Amount Bet | Amount Won/Lost | Model Predicted Favorite Not Cover | Model Predicted Favorite Cover | Not Cover Correct | Cover Correct\n--- | --- | --- | --- | --- | --- | --- | --- | --- \n2014 | 13 | 77% | \\\\$1,430 | \\\\$670 | 92% | 8% | 75% | 100%\n2015 | 14 | 64% | \\\\$1,540 | \\\\$350 | 86% | 14% | 58% | 100%\n2016 | 8 | 62% | \\\\$880 | \\\\$170 | 88% | 12% | 57% | 100%\n2017 | 17 | 53% | \\\\$1,870 | \\\\$20 | 76% | 24% | 38% | 100%\n2018 | 4 | 75% | \\\\$440 | \\\\$190 | 50% | 50% | 50% | 100%\n2019 | 6 | 67% | \\\\$660 | \\\\$180 | 83% | 17% | 80% | 0%\n**Total** | **62** | **65%** | **\\\\$6,820** | **\\\\$1,580** | **82%** | **18%** | **59%** | **91%**"},{"metadata":{},"cell_type":"markdown","source":"**Additional Considerations**\n\nUsing tournaments from 2010-2013, I determined the optimal betting strategy was to bet on games where the model's prediction varied by more than 3 points from the spread. But does this strategy hold moving forward? Using tournaments from 2014-2019, the model generated hypothetical winnings at the thresholds previously tested from 2010-2013.\n\nAs shown below, the optimal threshold was still 3. Even at a threshold of 1, the model had positive returns. This suggests the model generates more accurate predictions with additional training data. As the threshold increases 0 to 5, the model became more and more accurate. The model picked 50% of games correctly at a threshold of 0 and 78% correctly at a threshold of 5. However, the number of games selected for betting also decreases as the threshold increases. When the threshold increases from 3 to 4, the decrease in the number of games bet on outweighs the incremental increase in the percentage of games picks correctly. Therefore, the optimal threshold remains 3.\n\nMy analysis suggests an above average machine learning can beat the spread. The countdown to 2021 March Madness begins now. "},{"metadata":{},"cell_type":"markdown","source":"Model Difference from Spread | Amount Won/Lost | Games Bet On | Model Correct\n--- | --- | --- | --- \n0 | -\\\\$1,850 | 391 | 50%\n1 | \\\\$310 | 232 | 53%\n2 | \\\\$670 | 118 | 55%\n**3** | **\\\\$1,580** | **62** | **65%**\n4 | \\\\$1,420 | 31 | 74%\n5 | \\\\$960 | 18 | 78%\n"},{"metadata":{},"cell_type":"markdown","source":"The supporting code and sources are below.\n\nThis analysis is hypthetical and for research purposes. The Kaggle rules for this competition state \"You will not: (a) use or access the NCAA® Data for any commercial, gambling, or illegal purpose.\""},{"metadata":{},"cell_type":"markdown","source":"# Sources"},{"metadata":{},"cell_type":"markdown","source":"(in addition to Kaggle data)\n\nKen Pom:\n    \n2019\thttps://web.archive.org/web/20190317211809/https://kenpom.com/\n\n2018\thttps://web.archive.org/web/20180311122559/https://kenpom.com/\n\n2017\thttps://web.archive.org/web/20170312131016/http://kenpom.com/\n\n2016\thttps://web.archive.org/web/20160314134726/http://kenpom.com/\n\n2015\thttps://web.archive.org/web/20150316212936/http://kenpom.com/\n\n2014\thttps://web.archive.org/web/20140318100454/http://kenpom.com/\n\n2013\thttps://web.archive.org/web/20130318221134/http://kenpom.com/\n\n2012\thttps://web.archive.org/web/20120311165019/http://kenpom.com/\n\n2011\thttps://web.archive.org/web/20110311233233/http://www.kenpom.com/\n\n2010\thttps://web.archive.org/web/20100304023540/http://kenpom.com/rate.php\n\n2009\thttps://web.archive.org/web/20090315085050/http://kenpom.com/rate.php\n\nPrior to 2009 (only used for training, not validation): Downloaded directly from Ken Pom website (https://kenpom.com/index.php) as Wayback data are thin.\n\nSpread data:\n\nhttp://www.thepredictiontracker.com/ncaaresults.php \n\n(e.g., http://www.thepredictiontracker.com/ncaabb18.csv)"},{"metadata":{},"cell_type":"markdown","source":"# Supporting code and details"},{"metadata":{"trusted":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\n\nimport xgboost as xgb\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.metrics import mean_absolute_error\n\nimport sys\nif not sys.warnoptions:\n    import warnings\n    warnings.simplefilter(\"ignore\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Tournament games**"},{"metadata":{"trusted":true},"cell_type":"code","source":"tourney_results = pd.read_csv('../input/march-madness-analytics-2020/2020DataFiles/2020DataFiles/2020-Mens-Data/MDataFiles_Stage1/MNCAATourneyDetailedResults.csv')\ntourney_results = tourney_results[['Season','WTeamID','WScore','LTeamID','LScore']]\ntourney_results","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Regular season stats**"},{"metadata":{"trusted":true},"cell_type":"code","source":"regular_results = pd.read_csv('../input/march-madness-analytics-2020/2020DataFiles/2020DataFiles/2020-Mens-Data/MDataFiles_Stage1/MRegularSeasonDetailedResults.csv')\nregular_results = regular_results[['Season', 'DayNum', 'LTeamID', 'LScore', 'WTeamID', 'WScore']].copy()\nregular_results_swap = regular_results\nregular_results","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"regular_results.columns = [x.replace('W','T1_').replace('L','T2_') for x in list(regular_results.columns)]\nregular_results_swap.columns = [x.replace('L','T1_').replace('W','T2_') for x in list(regular_results.columns)]\nregular_data = pd.concat([regular_results, regular_results_swap]).sort_index().reset_index(drop = True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# data are duplicated, so only need to take T1_TeamID average, T2 will be average points allowed\nseason_statistics = regular_data.groupby(['Season','T1_TeamID'])['T1_Score','T2_Score'].agg(np.mean)\nseason_statistics.columns = [''.join(col).strip() for col in season_statistics.columns.values]\nseason_statistics = season_statistics.reset_index()\nseason_statistics.rename(columns={'T1_TeamID':'TeamID','T1_Score':'AvgPtsScored','T2_Score':'AvgPtsAllowed'}, inplace=True)\nseason_statistics","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Ken Pom data**"},{"metadata":{"trusted":true},"cell_type":"code","source":"kp_df=[]\nyears_list = [2019,2018,2017,2016,2015,2014,2013,2012,2011,2010,2009,2008,2007,2006,2005,2004,2003,2002]\nfor year in years_list:\n    temp_kp_df = pd.read_csv(\"../input/kenpomeffiencydata/KP\" + str(year) + \".csv\") \n    temp_kp_df = temp_kp_df[['team','conf','adjem','adjo','adjd','luck']]\n    temp_kp_df['Season'] = year\n    year_last = year   \n    if year==2019:\n        kp_df = temp_kp_df\n    else:\n        kp_df = kp_df.append(temp_kp_df)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"kp_df['team'] = kp_df['team'].str.lower()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# these teams are in the tournament historically and don't match to the team names file without these fixes\nkp_df.loc[kp_df['team']==\"st. louis\",\"team\"] = \"saint louis\"\nkp_df.loc[kp_df['team']==\"cal st. bakersfield\",\"team\"] = \"csu bakersfield\"\nkp_df.loc[kp_df['team']==\"illinois chicago\",\"team\"] = \"illinois-chicago\"\nkp_df.loc[kp_df['team']==\"texas a&m corpus chris\",\"team\"] = \"a&m-corpus chris\"\nkp_df.loc[kp_df['team']==\"texas a&m corpus chri\",\"team\"] = \"a&m-corpus chris\"\nkp_df.loc[kp_df['team']==\"nevada las vegas\",\"team\"] = \"unlv\"\nkp_df.loc[kp_df['team']==\"arkansas pine bluff\",\"team\"] = \"ark pine bluff\"\nkp_df.loc[kp_df['team']==\"mississippi valley st.\",\"team\"] = \"miss valley st.\"\nkp_df.loc[kp_df['team']==\"arkansas little rock\",\"team\"] = \"ark little rock\"\nkp_df.loc[kp_df['team']==\"texas el paso\",\"team\"] = \"texas-el paso\"\nkp_df.loc[kp_df['team']==\"wisconsin green bay\",\"team\"] = \"wisconsin-green bay\"\nkp_df.loc[kp_df['team']==\"wisconsin milwaukee\",\"team\"] = \"wisconsin-milwaukee\"\nkp_df.loc[kp_df['team']==\"md baltimore county\",\"team\"] = \"maryland-baltimore county\"\nkp_df.loc[kp_df['team']==\"winston salem st.\",\"team\"] = \"winston-salem-state\"\nkp_df.loc[kp_df['team']==\"southwest missouri st.\",\"team\"] = \"southwest missouri state\"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"teams_df = pd.read_csv('../input/march-madness-analytics-2020/2020DataFiles/2020DataFiles/2020-Mens-Data/MDataFiles_Stage1/MTeamSpellings.csv', sep='\\,', engine='python')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"kp_df = pd.merge(kp_df, teams_df, left_on=['team'], right_on = ['TeamNameSpelling'], how='left')\nkp_df = kp_df.drop(['TeamNameSpelling'], axis=1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# these teams haven't played in MM, fine that they don't match\ntemp = kp_df[kp_df['TeamID'].isna()]\ntemp['team'].value_counts(dropna=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"kp_df['Season'].value_counts(dropna=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"kp_df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Spread data**"},{"metadata":{"trusted":true},"cell_type":"code","source":"spread_df=[]\nyears_list = [18,17,16,15,14,13,12,11,10,'09','08','07','06','05','04','03']\nyear_int = 2019\nfor year in years_list:\n    #Note: 18 refers to 2019 and so on\n    temp_spread_df = pd.read_csv(\"../input/spread-data-03-to-18/ncaabb\" + str(year) + \".csv\")\n    temp_spread_df['Season'] = year_int\n    year_last = year\n    year_int = year_int-1    \n    if year==18:\n        spread_df = temp_spread_df\n    else:\n        spread_df = spread_df.append(temp_spread_df)\n    \nspread_df = spread_df[['Season','line','home','hscore','road','rscore']]\n\nspread_df = spread_df[(spread_df['hscore'] != \".\") & (spread_df['rscore'] != \".\") & (spread_df['line'] != \".\")]\nspread_df = spread_df.dropna()\n\nspread_df['rscore'] = spread_df['rscore'].astype(float)\nspread_df['hscore'] = spread_df['hscore'].astype(float)\nspread_df['line'] = spread_df['line'].astype(float)\n\nspread_df['rscore'] = spread_df['rscore'].astype(int)\nspread_df['hscore'] = spread_df['hscore'].astype(int)\nspread_df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# make teams W and L consistent with tourney data\nhome_team_wins = spread_df[spread_df['hscore']>spread_df['rscore']]\nhome_team_wins.rename(columns={'home':'WTeam','hscore':'WScore','road':'LTeam','rscore':'LScore'}, inplace=True)                          \n\naway_team_wins = spread_df[spread_df['hscore']<spread_df['rscore']]\naway_team_wins.rename(columns={'home':'LTeam','hscore':'LScore','road':'WTeam','rscore':'WScore'}, inplace=True)  \n#line (aka spread) is in terms of home team, multiply by -1 if the away team wins\naway_team_wins['line'] = away_team_wins['line']*-1\n\nspread_df = home_team_wins.append(away_team_wins)\nspread_df.rename(columns={'line':'WSpread'}, inplace=True)   \nspread_df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"spread_df['WTeam'] = spread_df['WTeam'].str.lower()\nspread_df['LTeam'] = spread_df['LTeam'].str.lower()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"team_list = ['WTeam','LTeam']\nfor team in team_list:\n    spread_df = pd.merge(spread_df, teams_df, left_on=[team], right_on = ['TeamNameSpelling'], how='left')\n    spread_df.rename(columns={'TeamID': team+'ID'}, inplace=True)\n    spread_df = spread_df.drop(['TeamNameSpelling'], axis=1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"spread_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"spread_df['Season'].value_counts(dropna=False)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Rankings data**"},{"metadata":{"trusted":true},"cell_type":"code","source":"rank_df = pd.read_csv('../input/march-madness-analytics-2020/2020DataFiles/2020DataFiles/2020-Mens-Data/MDataFiles_Stage1/MMasseyOrdinals.csv')\n# 133 is last day of regular season\nrank_df = rank_df.loc[rank_df['RankingDayNum'] == 133]\nrank_end_df = rank_df.groupby(['Season','TeamID'])['OrdinalRank'].agg(pd.np.mean).reset_index()\nrank_end_df.rename(columns={'OrdinalRank': 'Rank'}, inplace=True)\nrank_end_df","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Combine data"},{"metadata":{"trusted":true},"cell_type":"code","source":"tourney_results.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"spread_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"season_statistics.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"kp_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"rank_end_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"combined_data = pd.merge(tourney_results, spread_df, on = ['Season','WTeamID','WScore','LTeamID','LScore'], how = 'left')\ncombined_data.tail(5)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"reg_W = season_statistics[['Season','TeamID','AvgPtsScored','AvgPtsAllowed']].copy()\nreg_L = season_statistics[['Season','TeamID','AvgPtsScored','AvgPtsAllowed']].copy()\nreg_W.columns = ['Season','WTeamID','WAvgPtsScored','WAvgPtsAllowed']\nreg_L.columns = ['Season','LTeamID','LAvgPtsScored','LAvgPtsAllowed']\n\ncombined_data = pd.merge(combined_data, reg_W, on = ['Season', 'WTeamID'], how = 'left')\ncombined_data = pd.merge(combined_data, reg_L, on = ['Season', 'LTeamID'], how = 'left')\ncombined_data","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"kp_W = kp_df[['Season','TeamID','adjem']].copy()\nkp_L = kp_df[['Season','TeamID','adjem']].copy()\nkp_W.columns = ['Season','WTeamID','Wadjem']\nkp_L.columns = ['Season','LTeamID','Ladjem']\n\ncombined_data = pd.merge(combined_data, kp_W, on = ['Season', 'WTeamID'], how = 'left')\ncombined_data = pd.merge(combined_data, kp_L, on = ['Season', 'LTeamID'], how = 'left')\ncombined_data","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"rank_W = rank_end_df[['Season','TeamID','Rank']].copy()\nrank_L = rank_end_df[['Season','TeamID','Rank']].copy()\nrank_W.columns = ['Season','WTeamID','WRank']\nrank_L.columns = ['Season','LTeamID','LRank']\n\ncombined_data = pd.merge(combined_data, rank_W, on = ['Season', 'WTeamID'], how = 'left')\ncombined_data = pd.merge(combined_data, rank_L, on = ['Season', 'LTeamID'], how = 'left')\ncombined_data","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Having features relative to winning team would cause a data leak since the features incorporate the outcome**\n\nWe instead make the variables relevant the favorite team. If WSpread is positive, it means the favorite won. If WSpread is negative, the underdog won."},{"metadata":{"trusted":true},"cell_type":"code","source":"#we'll also include 0 in this one:\nfavorite_won = combined_data[combined_data['WSpread']>-0.01]\nfavorite_won","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"favorite_won.columns = [x.replace('W','F').replace('L','U') for x in list(favorite_won.columns)]\nfavorite_won","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"underdog_won = combined_data[combined_data['WSpread']<0]\nunderdog_won.columns = [x.replace('W','U').replace('L','F') for x in list(underdog_won.columns)]\n#The winning spread is related to the underdog, swap by -1 to make relative to favorite\nunderdog_won['FSpread'] = underdog_won['USpread']*-1\nunderdog_won = underdog_won.drop(['USpread'], axis=1)\nunderdog_won","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"combined_data = favorite_won.append(underdog_won)\ncombined_data","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"combined_data['Season'].value_counts(dropna=False).sort_index()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"combined_data['FAvgPointMargin'] = combined_data['FAvgPtsScored']-combined_data['FAvgPtsAllowed']\ncombined_data['UAvgPointMargin'] = combined_data['UAvgPtsScored']-combined_data['UAvgPtsAllowed']","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# make variables we will use as features differences (cuts down the number of features to potentially generalize better)\nvarlist = [\n'AvgPointMargin',\n'adjem',\n'Rank'\n]\nfor var in varlist:\n    combined_data[var+'Diff'] = combined_data['F'+var]-combined_data['U'+var]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# add target variables:\ncombined_data['ScoreDiff'] = combined_data['FScore']-combined_data['UScore']\n\ncombined_data['FavoriteCovered'] = 0\ncombined_data.loc[combined_data['FSpread']<combined_data['ScoreDiff'], 'FavoriteCovered'] =  1\n\ncombined_data['FWon'] = 0\ncombined_data.loc[combined_data['FScore']>combined_data['UScore'], 'FWon'] = 1","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Train model to predict wins and compare to historical Kaggle leaderboard"},{"metadata":{"trusted":true},"cell_type":"code","source":"features = [\n'AvgPointMarginDiff',\n'adjemDiff',\n'RankDiff',\n'FSpread'\n]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def predict(df,season,features_array,dep_var):\n    # train on prior seasons, validate on current season\n    train_df = df[df['Season']<season]\n    valid_df = df[df['Season']==season]\n\n    X_train = train_df[features_array].reset_index(drop=True)\n    y_train = train_df[dep_var].reset_index(drop=True)\n\n    X_valid = valid_df[features_array].reset_index(drop=True)\n    y_valid = valid_df[dep_var].reset_index(drop=True)\n    \n    X_train_xgb = xgb.DMatrix(X_train, label = y_train)\n    X_valid_xgb = xgb.DMatrix(X_valid)\n\n    param = {'max_depth':3,'eta':.001,'seed':201,'objective':'binary:logistic',\n             'eval_metric':'mae', 'gamma':0, 'nthread':-1}\n\n    num_round=2000\n    \n    xgb_train = xgb.train(param, X_train_xgb, num_round)   \n    pred = xgb_train.predict(X_valid_xgb)\n        \n    return pred, y_valid, valid_df, xgb, xgb_train","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#this is the scoring metric for Kaggle competitions\ndef LogLoss(predictions, realizations):\n    predictions_use = predictions.clip(0)\n    realizations_use = realizations.clip(0)\n    LogLoss = -np.mean( (realizations_use * np.log(predictions_use)) + \n                        (1 - realizations_use) * np.log(1 - predictions_use) )\n    return LogLoss","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#evaluate log loss during the last 4 seasons\nlosses = 0\nseasons = [2016,2017,2018,2019]\nfor season in seasons:\n    pred, y_valid, valid_df, xgb, xgb_train = predict(combined_data,season,features,'FWon')    \n    loss = LogLoss(pred, y_valid)\n        \n    print(\"Season\",season,\"valid:\",loss)\n    losses = losses+loss\n    \nlosses_avg = losses/4\nprint(\"All seasons :\",losses_avg)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Average (2016-2019): 19th percentile\n\n2016: 1st percentile\n\n2017: 29th percentile\n\n2018: 31st percentile\n\n2019: 14th percentile\n\n\nSource: Kaggle leaderboards\n\nBackup:\n\n2016: 9/597\n\n2017: 126/441\n\n2018: 292/933\n\n2019: 122/866"},{"metadata":{},"cell_type":"markdown","source":"# Train model to predict point difference"},{"metadata":{},"cell_type":"markdown","source":"We will train the model. We will then determine an optimal betting strategy during 2010-2013. We will then simulate this optimal betting strategy during 2014-2019. "},{"metadata":{"trusted":true},"cell_type":"code","source":"combined_data","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"From: https://www.thelines.com/betting/point-spread/\n\n\"Point spreads are usually set with -110 odds, but pricing often fluctuates at online sportsbooks. This is the sportsbook operators’ house edge. The odds guarantee the sportsbook operator will see a little money over time. When the odds are set at -110, the bettor must wager 110 to win 100 (or 11 to win 10).\"\n\nHere, we will always place 110 bet and  win 100"},{"metadata":{"trusted":true},"cell_type":"code","source":"def predict(df,season,features_array,dep_var):\n    train_df = df[df['Season']<season]\n    valid_df = df[df['Season']==season]\n\n    X_train = train_df[features_array].reset_index(drop=True)\n    y_train = train_df[dep_var].reset_index(drop=True)\n\n    X_valid = valid_df[features_array].reset_index(drop=True)\n    y_valid = valid_df[dep_var].reset_index(drop=True)\n    \n    X_train_xgb = xgb.DMatrix(X_train, label = y_train)\n    X_valid_xgb = xgb.DMatrix(X_valid)\n\n    param = {'max_depth':3,'eta':.01,'seed':201,'objective':'reg:squarederror',\n             'eval_metric':'mae', 'gamma':1, 'nthread':-1}\n\n    num_round=200\n            \n    xgb_train = xgb.train(param, X_train_xgb, num_round)   \n    pred = xgb_train.predict(X_valid_xgb)\n        \n    return pred, y_valid, valid_df, xgb, xgb_train","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#2009 is the first season we have without a data leak in Ken Pom. \n#We still include the training data prior to 2009 that has a leak as the validation performance is better with more data, even if some of the years of training data have Ken Pom leaks\n#We use all prior years and validate on each season since 2009\n\nseasons = [2009,2010,2011,2012,2013,2014,2015,2016,2017,2018,2019]\nfor season in seasons:\n    # train model using all seasons in the past, validate on current season  \n    pred, y_valid, valid_df, xgb, xgb_train = predict(combined_data,season,features,'ScoreDiff')  \n    # mean average error is average difference between predicted and actual score difference\n    mae = mean_absolute_error(pred, y_valid)\n    print(season,'MAE',mae)\n    \n    valid_df['PredSpread'] = pd.array(pred)\n\n    valid_df['PredFavoriteCovered'] = 0\n    valid_df.loc[valid_df['PredSpread']>valid_df['FSpread'],'PredFavoriteCovered'] = 1     \n    \n    valid_df['DollarsBet'] = 110\n    #wrong: lose $110\n    valid_df['DollarsWonLost'] = -110\n    #right: gain $100\n    valid_df.loc[(valid_df['FavoriteCovered']==1) & (valid_df['PredFavoriteCovered']==1),'DollarsWonLost'] = 100   \n    valid_df.loc[(valid_df['FavoriteCovered']==0) & (valid_df['PredFavoriteCovered']==0),'DollarsWonLost'] = 100 \n\n    # We exclude 2009 from the simulation because the MAE is high. We don't have enough training data.\n    if season==2010:\n        all_validation_bets = valid_df\n    if season>2010:\n        all_validation_bets = all_validation_bets.append(valid_df)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# this is the feature importance from last iteration above\nxgb.plot_importance(xgb_train)\nplt.rcParams['figure.figsize'] = [20, 20]\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"all_validation_bets.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# What is optimal threshold to cutoff in difference? \n# Use 2010-2013 to determine threshold and then test strategy in 2014-2019\ntrain = all_validation_bets[all_validation_bets['Season']<2014]\ntest = all_validation_bets[all_validation_bets['Season']>2013]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"In the 10 games below, we can see 2 differ by more than 3 points. In a strategy where we set the threshold at 3, we would bet on these two games and not bet on the other 8. More specifically, we would bet on San Diego State v Florida Gulf Coast and UMiami v Marquette. If we set the threshold at 1, we would bet on 6 out of 10 of the games."},{"metadata":{"trusted":true},"cell_type":"code","source":"train['Difference'] = abs(train['PredSpread']-train['FSpread'])\ntrain[['Season','FTeam','FScore','UTeam','UScore','ScoreDiff','FSpread','PredSpread','Difference']].tail(10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# keep games where we differ from spread by at least 3\ntrain0 = train[abs(train['PredSpread']-train['FSpread'])>0]\ntrain1 = train[abs(train['PredSpread']-train['FSpread'])>1]\ntrain2 = train[abs(train['PredSpread']-train['FSpread'])>2]\ntrain3 = train[abs(train['PredSpread']-train['FSpread'])>3]\ntrain4 = train[abs(train['PredSpread']-train['FSpread'])>4]\ntrain5 = train[abs(train['PredSpread']-train['FSpread'])>5]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The optimal threshold is 3"},{"metadata":{"trusted":true},"cell_type":"code","source":"for i in range(0,6):\n    train_temp = train[abs(train['PredSpread']-train['FSpread'])>i]\n    train_temp['DifferenceVspread'] = i\n    train_temp['GamesBetOn'] = 1\n    if i==0:\n        appended = train_temp\n    else:\n        appended = appended.append(train_temp)\n        \nsummary_df = appended.groupby(['DifferenceVspread'])['DollarsWonLost','GamesBetOn'].agg(np.sum).reset_index()\nsummary_df = summary_df.style.format({\n    'DollarsWonLost': '$ {:,.0f}'.format\n})\nsummary_df","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Does this same strategy hold into the future? Let's see the optimal threshold if we include all seasons. This same strategy does hold into the future."},{"metadata":{"trusted":true},"cell_type":"code","source":"for i in range(0,6):\n    test_temp = test[abs(test['PredSpread']-test['FSpread'])>i]\n    test_temp['DifferenceVspread'] = i\n    test_temp['GamesBetOn'] = 1\n    test_temp['ModelCorrect'] = test_temp['DollarsWonLost']==100\n    if i==0:\n        appended = test_temp\n    else:\n        appended = appended.append(test_temp)\n        \nsummary_df = appended.groupby(['DifferenceVspread'])['DollarsWonLost','GamesBetOn','ModelCorrect'].agg(np.sum).reset_index()\nsummary_df['ModelCorrect'] = summary_df['ModelCorrect'] / summary_df['GamesBetOn']\nsummary_df = summary_df.style.format({\n    'DollarsWonLost': '$ {:,.0f}'.format,\n    'ModelCorrect': '{:,.0%}'.format\n})\nsummary_df","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"How many games do we bet on in the test set (2014-2019)?"},{"metadata":{"trusted":true},"cell_type":"code","source":"test.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test = test[abs(test['PredSpread']-test['FSpread'])>3]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# we bet on 62 out of 391 games\ntest.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"62/391","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's evaluate our model accuracy:"},{"metadata":{"trusted":true},"cell_type":"code","source":"test['ModelCorrect'] = 0\ntest.loc[test['DollarsWonLost']>0,'ModelCorrect'] = 1 \ntest['Seasons'] = \"2014-2019\"\nsummary_df = test.groupby(['Seasons'])['ModelCorrect'].agg(np.mean).reset_index()\nsummary_df = summary_df.style.format({\n    'ModelCorrect': '{:,.0%}'.format\n})\nsummary_df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test['ModelCorrect'] = 0\ntest.loc[test['DollarsWonLost']>0,'ModelCorrect'] = 1 \n\nsummary_df = test.groupby(['Season'])['ModelCorrect'].agg(np.mean).reset_index()\nsummary_df = summary_df.style.format({\n    'ModelCorrect': '{:,.0%}'.format\n})\nsummary_df","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's look at a few examples of how the above works to unpack this a bit. \n\n(721) Louisville was the favorite against Manhattan and won 71-64, a difference of 7 points. Louisville was favored to win by 17 points, as represented by the spread. The model predicted Louisville would win by 11.7. Therefore, the model predicted Louisville would not cover. Because Louisville did not cover the spread, the model won \\\\$100 off \\$110 bet (i.e., bet \\\\$110, returned \\$210).\n\n(732) Wisconsin was the favorite against American and won 75-35, a difference of 40 points. Wisconsin was favored to win by 14.5 points, as represented by the spread. The model predicted Wisconsin would win by 11.4 points. Therefore, the model predicted Wisconsin would not cover. Wisconsin won by far more than 14.5. Therefore, the model lost the full \\\\$110 bet (i.e., bet \\$110, returned \\\\$0)."},{"metadata":{"trusted":true},"cell_type":"code","source":"test[['Season','FTeam','FScore','UTeam','UScore','ScoreDiff','FavoriteCovered','FSpread','PredFavoriteCovered','PredSpread','DollarsBet','DollarsWonLost']].head(10)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Evaluate bets"},{"metadata":{"trusted":true},"cell_type":"code","source":"test['GamesBetOn'] = 1\n\ntest['NotCoverCorrect'] = np.where((test['PredFavoriteCovered']==0) & (test['FavoriteCovered']==0), 1, 0)\ntest['CoverCorrect'] = np.where((test['PredFavoriteCovered']==1) & (test['FavoriteCovered']==1), 1, 0)\n\ntest_all = test.copy()\ntest_all['Season'] = 'Total'\ntest_appended = test.append(test_all)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"summary_df = test_appended.groupby(['Season'])['GamesBetOn','ModelCorrect','DollarsBet','DollarsWonLost','PredFavoriteCovered','NotCoverCorrect','CoverCorrect'].agg(np.sum).reset_index()\n\nsummary_df['ModelCorrect'] = summary_df['ModelCorrect']/summary_df['GamesBetOn']\nsummary_df['TookCover'] = summary_df['PredFavoriteCovered']/summary_df['GamesBetOn']\nsummary_df['TookNotCover'] = 1-summary_df['TookCover'] \n\nsummary_df['CoverCorrect'] = summary_df['CoverCorrect']/summary_df['PredFavoriteCovered']\nsummary_df['NotCoverCorrect'] = summary_df['NotCoverCorrect']/(summary_df['GamesBetOn']-summary_df['PredFavoriteCovered'])\n\nsummary_df = summary_df[['Season','GamesBetOn','ModelCorrect','DollarsBet','DollarsWonLost','TookNotCover','TookCover','NotCoverCorrect','CoverCorrect']]\n\nsummary_df = summary_df.style.format({\n    'ModelCorrect': '{:,.0%}'.format,\n    'TookCover': '{:,.0%}'.format,\n    'TookNotCover': '{:,.0%}'.format,\n    'CoverCorrect': '{:,.0%}'.format,\n    'NotCoverCorrect': '{:,.0%}'.format,\n    'DollarsBet': '$ {:,.0f}'.format,\n    'DollarsWonLost': '$ {:,.0f}'.format\n})\n\nsummary_df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"summary_df = test_appended.groupby(['Season'])['GamesBetOn','ModelCorrect','DollarsBet','DollarsWonLost','PredFavoriteCovered','NotCoverCorrect','CoverCorrect'].agg(np.sum).reset_index()\n\nsummary_df['ModelCorrect'] = summary_df['ModelCorrect']\nsummary_df['TookCover'] = summary_df['PredFavoriteCovered']\nsummary_df['TookNotCover'] = summary_df['GamesBetOn']-summary_df['TookCover'] \n\nsummary_df = summary_df[['Season','GamesBetOn','ModelCorrect','TookNotCover','TookCover','NotCoverCorrect','CoverCorrect']]\n\nsummary_df = summary_df.style.format({\n    'DollarsBet': '$ {:,.0f}'.format,\n    'DollarsWonLost': '$ {:,.0f}'.format\n})\n\nsummary_df","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Game-level detail for bets placed"},{"metadata":{"trusted":true},"cell_type":"code","source":"pd.options.display.max_columns = None\npd.options.display.max_rows = None","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test[['Season','FTeam','FScore','UTeam','UScore','ScoreDiff','FSpread','PredSpread','PredFavoriteCovered','DollarsWonLost']]","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}