---
title: 'Right, Left, Shoot: March Madness EDA'
date: '`r Sys.Date()`'
runtime: shiny
output:
  html_document:
    fig_caption: true
    toc: true
    fig_width: 7
    fig_height: 4.5
    theme: paper
    highlight: tango
    code_folding: hide
---

```{r setup, include=FALSE}
knitr::opts_chunk$set(echo = TRUE)
```


<center><img src="data:image/jpeg;base64,/9j/4AAQSkZJRgABAQAAAQABAAD/2wCEAAkGBxMRERUSEBMWFRUXGBoYFxcVGBgeFxgYFxkYGBcXGhcbHSggGBslGxcXIjEhJSkrLi4uGB8zODMsNygtLisBCgoKDQ0OFhAQFy0lHyYtNS0rMDc3Ky0vMi03LTcrLy0xNjYuMC03LystKzUtKy0uLy0rLS03LSsrLS0tLS0tK//AABEIAKsBJgMBIgACEQEDEQH/xAAcAAACAwEBAQEAAAAAAAAAAAAAAwECBAUGBwj/xABDEAABAwEEBQgFCwMFAQEAAAABAAIRAwQSITETIkFRUgUUMmGRktHSBmJxgaIVI0JTcoKhseHi8FSTsgckM2PB8UT/xAAYAQEAAwEAAAAAAAAAAAAAAAAAAQIDBP/EACkRAQABAwMCBgEFAAAAAAAAAAABAgMRE6HwIVIEEjFRkeHRIjJBgbH/2gAMAwEAAhEDEQA/AN77bWcXE2ivJc7KtUA6RyAdAHUFHOav9RaP79XzL19p9A6YD3Nr141nBjRSJ2uuiWY7hJXE+RaN5zTVtQLZmRZhiATdG2dkdYXDNm/MzireXBVY8RMzireXL5zV/qLR/fq+ZHOav9RaP79XzLrWXkClUc1ralsl2erZ9USBedAOGOYnIrsP9AWAE85rmBMAUZPUNTNRoeI7t5V0PEd28vI85q/1Fo/v1fMr0a1Zzmt5xaNYgYV6u0xxrov5Eot6dS2NPC5tmDs4GBG1VHI1GboqWwnVgAWaTeAMDrE/gU0b/dvKdDxHdvLo1ORKgMc6te7/AJKuqYqGXG/0fmxiNroVKPI9RxxtVrEtY4TUrfTYXGTejAiFss3oM17Q42i0tkdF4pBwxyIuLLyv6L07OAXWi0uvcOgwAjO8BhittOvky107nJkjk7kupVpsfzu0C8BI01WReq6MRr4ggO9460DkqsWPe20WsEAFjH1agccAXSL+wERCz/JFANvaa1DLV/2172xuBETvXdHoCzPnVf2xS8iiLVzHX/ZRpXcfcuVX5MqNrCmLTanNLHvBbWqEk07wLQA+M2/iEx/I9WbrbVarxc9rZq1YJawPA6eZkj2tK12v0Jp0mF7rVX1QTA0IJPCJaBJXLbyPQOVotRwn/wDPmdg9bPsKadzkyad3ky1UuSXuIAtVsg6SDpKsHR1GsGTsJvTPUinyPUJANptmIMxUqThVFLDXyxveyFrd6EMuX2Wqu8AS25oSXYZN1YM+1YD6NAAH/f7YAZQke6MFOnc5Mp07vJkvlLk6pSp6QWu0PxY0gVqsy68Thf3AR7TuVvkyqatak202qWAXL1arruOQOvgCcJSavItBpANW1zBLm3bPLLnECMME8+jQz/3+O25Z/BRp3c/co0rvJloPIr8ItdqIL2NkVauTwwz0vWShyW7SCnzy1ElmkF2pWILbzmjJxjIHHeVy+WOS6tlGkOlNDK/UAa5hGHzgbgGmAWuwzgxAnC07Z/FY3LlVFWJifmWNy7VROJifmXes/JlVza55zappOe0AVauNxhdjr7YjNam8g1b13ndqOAyq1TrXi1wkPgRAOMZ7F5cGNqgvMhrQ5z3GGsbJc8nYBt6zkMzCpTfz0xPzKlN+fTr8y9MzkapeA53aTLS7VrVTEMpuunXJvG+TlkBvSrNye592LVawXNe4g1amGjdcIweTn1KKPorUaAKxtOkIvO0DaRptnJgc7FzgIBO0zG5UfyBTZi+pbaYAJBe2gASATdBjMwunyXOTLp07vJlp+Q62AFqtJN94I01SbrdIGui/Ml1MD7wWGnYa5dQDq9qaKsSTVrarrzgW4uxOrl1ptn5BY/FjracJwbZ8sSMQOorTZvRUPc1pdbWyc3so3RgTJIB2iPaQk27nvvJNq7yZLHJVUtLharU0w4hlSrVDzFO+AAH44/gm1OQqskNtdpJgXfnqmLwXhzcH7Awn3q/KXonSoBpfaq5vGMNDIEYuMtGAwBjeFjs3INBzw3nFobJzdzaBnjgN+r7cE06+TKdO7yZOtPJZYJ57aTPQ+efLtWm4ANvySdJsyjFMHIlQuLedWsazmyalSNWoynib+Zv3h1Bbh6AM/qq3ZS8ix8q+iVOgwPfabQ6XAAN0MznOs0bk07nJk07vJkr5GqXSed2l2rfbFWrBB0hBm/AwYMCRmdyrX5JqMDzzu0uuiqdStVP/ABBhAdrYE3/wWRnJNC6Xaa1AQYH+21oIwA24G97F2rP6DMexr22qvDgHZUsnCdjE07nJk07vJlw+WqLqDrrLXXcZcHDT1JF2IJh+EycDjgudzmr/AFFo/v1fMvQ8o+iVOi4XqtqcCJL2toXW/aJaIWGnyPQIJ0tqERmLNtIH/oKzqs3pnpVj+5Z1Wb8zmKt5cznNX+otH9+r5kc5q/1Fo/v1fMu5yd6MU6zgGvtYBnXc2hdECYJAOOxb7V6CsY0uFe0Pj6LG0bx9ksCjRv8AdvKuh4ju3l5CpaLScKVasTBJm0VAABG94k4rRRZai2XVa8jOLS8jATgRUxwW13IVnc03qlpIGwiyycYkDdEmdwK1cm+i9Kq+4x9qaB9Its90YTBgGN2S6ZpqmzFGP1Z9cymLF/u3+mz/AE8tVR1eux9Wo9op03AVHudBLngkXiYwA7ELvejvoyyxve9tWpUL2tbr3IAaXHANaNrtu5CtbpmmmIq9XZapqpoiKvV3Up1lYTJY0nfdE9qahaNC6dBjTLWtBykAAwmIQgVUs7HGXMaTvIBKG2ZgMhjQd90SmoQCpVotd0mh0ZSAfzV0IEczp/Vs7o8E9Ch2SCKlMOEOAI3ESEvmlP6tndHguI60VL4AqDok3CTfMEawM9EZHDaMRtZpKv8AHO8EHcYwAQAABkBkpXC0lXf8TvBAq1Jgk+4nwQdl1nYc2N7BtzTF5+hWqFoN8OzxaXQcff7M9ivpKm/4neCDukTmvDcocl8nkyKFakTJOhBaN8lgN0EztGK7mkqb/id4LNyhyk+kycS6YAl0HHHEAxhjkmKZ/dGVaqfNDju9HLI2C82xwJiJaMnFubADjdMY7t4XovRuz2VheLNRLCALz3NN54Mx84SS7LeuK30hq6hLIBi9DqmqLzQT0NbAzl1LsurvJ1ScpxJSKbcelOFabcUz0x8O6qvph3SAPtEri6Srv/F3gjSVd/4u8EaOyyk1vRAHsACuuALRUvgF46JN2TfzGtn0dmW3PYrNrVDjPxO8EHZq0Gu6TWujKQD+arzSn9Wzuj27t65Okq7/AIneCNLV3/E7wQdxUqUmuEOAcOsA/mudyRVc46zw7A4tm7n7T7F1ECeZ0/q2d0bc9ia1oAAAgDIDIKUIIc0EQRI3FK5nT+rZ3R4JyEFadMNENAA3AQrIQgQLHT+rZ3R4JlOi1vRaGznAAnsV0IBCEIMpra7mio2RBLSMRIwxkTMFWvu+sp9h86C/WcL7TEapzbh/7mpnrZ/Pegi+76yn2HzqG1zJGD8AdQARnnLupWnrZ/PeobWgkQDgDqj259iC2mPA74fMjTHgd8PmU6b1XdiNN6ruxANrYgFrhO+P/CmpD9YgQ4Y55bDtCtoBvd3neKBqClaAb3d53il16d0SC6ZH0jvCCHcn0zm0H2qPk2lwBa0IMnybS4ApbyfTGTQPYtSEGT5NpcAR8m0uALWhBk+TaXAFmt1WzWcN0z2UmuN1pe660uOTZOEndtXUXxn/AFGfXqVrTTqVK0N1qNNo+YugC4XPvBgJeHA3scI3K9FHmnGVa6vLHpl9bo0KTxeYQ4b2ukdoKl3J9M5tn2r8+ULXXlgpUa1Ko0RqMrtdVPzgF4MaC5wDDhvY6cl7z/TT0itlqtZp1ajzSbSL3MqNEi8W6IteWh11wJzmbpMws4z/ADCtNyZnEw+jfJtLgCPk2lwBa0KWjK3k6kDIaAd4zR8m0uALUhBk+TaXAEfJtLgC1oQJoWVjOg2PYnJNqOAgkS5ow6yJU6Ab3d53igaoe6AScgJ7EvQDe7vO8UuvZpa4AumD9I7vagvpjwO+HzI0x4HfD5lbTeq7sUab1XdiCNMeB3w+ZLp1nETeYM8CMRjt1k7Teq7sSqdW8J1Rngc8/agnSHjZ2fvRpDxs7P3qZ62fz3onrZ/PegLHVDg4h4fDiNUQARALf5vQpsrpvawdDowjVwGr/N6EFb0vcJYSIN36TQcp9sGFa71N/nuVC2XuwpnL7Q9uB9ytovUZ/PuoJu9Tf57lAqQYI2A6oJ37gjReoz+fdQxwaYugYA6oJ37ggvphud3XeCNMNzu67wRph1913gjTDr7rvBACqCYxx3gj8wmJDzeLYkY5wdx3iFfRnjd8PlQMSLYJYR1t/wAgr6M8bvh8qh1GcC5xGB+jsM7kEc2G9/fd4o5uN7u+7xTkIEus4jN3fd4q1B0taTmQPyV3ZKlm6Dfsj8kDEIQgF8p9K6rqtSqWFtFlSv8AOB9Nri+lYqRdWvC9eAMvALcTcGIkFfU69UMa57jAaC4nqAkr4CbDaa1XnNmrVKlQ33tbSFQuZpnF72BzJYAb0GYyWlunKtVUQ9iw2kVA2/ZyW1CQSyoCNbXOD99O2HOMXDEFdb/TmxNbXttTVvtdQoG6TlRoMzaSbpF66RObCdq+W22lbLP83a9OGQ/VqPe1vzl4Ph3RcTffIk9N29fTf9K+Sa7RVtdfVFoDTTaTJLZc7SOMky69tMpcpqpx06e/RNFVFUTievt1/D36EIWaQlVmyQJIzyJH5Jqo7pN96CnNhxP77vFHNhxP77vFOQgyV6AF0y7pNzcSOkNkrWqVad4RliDhvBlRozxu+HyoGKHOgEnIYqmjPG74fKl16RLXazjgdjd32UF9ONzu67wRpxud3XeCnTDr7rvBGlHX3XeCA0w3O7rvBLpm8Jutx355+xN0o6+67wSaYvCQxuM555/ZQXu9TP57kXepn89yjReoz+fdRovUZ/PuoCyvkOhzXQ4jVybGbT1hSizZHodI4M2ZYHrQgW6lruOjYZjHC9hsdh7YU6P/AK29o8FV9LXcdGcY1mui9A2iRkjR+pU7/wC9BbR/9be0eCGVGtJDg1hgbRjn/Peq6P1Knf8A3q1FzWkyC0wOk6ZGOWsUDOdM4294I50zjb3graVvEO0I0reIdoQQyu0mA5pPUQmJFR0loa4Z7IOwq+jdxHsHggYhL0buI9g8FSteaJvTiNg2kBA9CEIIdkkWambjdY9EcO72J7skuzdBv2R+SCdGeI/D4I0Z4j8PgmIQKa0OaQ7WBLgZAxEkYjJXYwAQAANwyVaGR+07/IpiClWk1wLXAOBzBAIPuKmlTDQGtAAAgAZADIAKyECjJcRJGAyjaTvHUp0Z4nfD4IHTP2W/m5MQL0Z4nfD4KLsOEuJzzj/wJqo/Nvv/ACQXQhCAQlWhxAEbS0dpAU3HcX4BAxBKXcdxfgEq0sdcdrfROwbkDOcs4294I5yzjb3graVvEO0I0reIdoQV5yzjb3gksAcJFNpBnGRjjnktGlbxDtCzNAOLWPjHJ8DPYL6C+j/6m9o8EaP/AKm9o8FW56lTv/vRc9Sp3/3oGWRl0HVa3WJhnuxPrIRZGQDqXNYnOSZjWPX25IQUdS13OunGMQ7OBtEiNqm6eF/fHmVXUhfc646SALzXdIDqkREq131anePmQF08L++PMoY9ocb0tMDpOBwxyxKm76tTvHzIpva0mZaYHTdmMcpJ60F9Mzib2hGmZxN7Qrc4Zxt7QjnDONvaEAyq0nAieohMWepVDi0Mc2Z9uw7AQr3X8Te4fMgak2zoH2t/yCm6/ib3D5lV9J7hBc2JBwadhB4upBe47i/D9UXXcX4fqmIQLLHcX4fqrU2wANwA7FZCAQhCBdDI/ad/kVZ1QDMge9c23VrQ27zemx4JqXrxiDeF3GRA6WIDshgvK8ocj2hxqVn2WkHOvOfNZ93oFpddDom6SPZvKD3gqDeO0KQZyXzdnJxquhlKm9xdeEVKjNcEvlpvSBq4SBIkGJM9zkqjbqFPR07PSAvaodUJAF2B9IkCQ3aTicEHqR0z9lv5uTEmhe+nF6629dmJ1picYlOQCpUYTEGI6ldCBd13F+H6ouu4vw/VMQgzWhp1ZdOs3Z6wWlLrU7wgGDIOU5GVF1/E3unzIGqCd6XdfxN7p8yXaWvuO1m9E/RO77SBmmZvb2hGmZvb2hTzlnG3tCjnLONvaEBpmb29oSaeIlrXRjEOEZ+1O5wzjb2hJZdOLQ8jqcYz2ayC8Hhf3h5kQeF/eHmUR6tTvHzIu+rU7x8yC1kp3QdW7Lic5JnMndjOCFFjphocAwtlxOJkmc3ZmJKEBoml7nBxDsA4B2GAkYHAYFW0Xru7R4JbqRc90tpuECJ6QO2cD1Qp5uPq6f8APuoL6L13do8FanSAJMknr6v/AKlc3H1dP+fdRTs8EkAMwA1Yxic9XrQaIRCpozxH4fBGjPEfh8EDIQkVHFpbJJBMZTsJ2Cditzhvrdx3ggahK5w31u47wRzhvX72u29cIGoQhAm2OIYSFkosc5oOmiQDGO37y216d5pGUrj8q8j0zTJq1CxjdYuaS3IEYkbMcvYg3aF31387yNC764fz7y818n2K6086MYEONR24GL045jA9W9dbk3kBjBeZUc8Oa0C+S4QLxBF45m8ZOZgbkD6BeXlgflOIyMEbJ61r5q/609h8yVSsJza6IkYDr9vUqW17aLQ6rXuNJDQXE5uMAZ7Sg0c1f9aew+ZVfZ3gE6U5bj5liNupA3ecid0mTJgRjjit5sjvrD+Pigy2dznk/OXcBntmesfwp+id9cOz9yX8nYxIwA2b58FxLVyLZadW7UtD2vdeeGGo7EHAw2ctw35YoO/onfXDs/cq6zXN+cvSYj+Erg2Xkay1HltO0vLjeloqOwjAi7OrEbI29a9BQ5ODXB0jDqQbkIQgEKtSoGiT/JwGAS+ct9buP8EDkJPOW+t3H+CpWtQDSRekA/Qdu9iDShU0Z4j8PgjRniPw+CC6SKIGTiPePBXuHiPw+CQyzQMWMOeJzOO3VQN0fru7R4I0fru7R4KnNx9XT/n3Uc3H1dP+fdQWstNrbwaSdYkyZhxgkdWzBCiyAgEG4NYwGZAbJ9behBQtBe7UYcsQRey+lhh1dSnRD6ofChzTfdNNpbAhwOsTjMgjCMIxPuU6MfV/kgjRD6ofCinRMkgXBAyjHPqU6MfV/khlMySBdEDcZz3FBfRu4z2N8EaN3Gexvgpuu4h3f1RddxDu/qgBSMglxMdQ9mwdaYkPeWkXnCCYyjYTv6lfTt4ggYk2sw33t/yCtp28QSrRVaWwCCZb/kEDTXbv/Ao07d/4FMQgXp27/wA1FeiyqwseA5jhBByIKY7JUs3Qb9kfkgwjkCzQ0aJsNi6MYECAQJzyxzwG4LdZqDabGsYA1rQGtAyAGACYhBSjkfa78yk2+wUq7QyswPaHBwB2OaZafcQm0Mj9p3+RTEHNqcg2dxJdSaZEYzlM3YmLs7MoJGS6SEIKDpH2D83LNauSqNSoKlSmHPDS0OxkA44e/bmFpHSPsH5uV0GGzcj0Kb77KYDgXGcc3dI4nM7+s7ytj6gGZVlR3SHvQRp27/wKNO3f+BTEIM1oqg3QD9Ju/iC0pNqyH2m/mFbTt4ggYq1GyCDtEdqrp28QVK1paGuIcJAJHuCC2jdxnsb4I0buM9jfBTddxDs/VF13EOz9UEaN3GexvgksowMWBxxxwxxT7ruIdn6pNOlA1mScccMcUE6L/qH4I0X/AFD8FNwfV/l4ouD6v8vFBNkHS1Wt1jg0jcMXQOl+iEWUEA3mtbrHBpmRsJwGKlAtzTfd82YgazXRePWJERgrXfVf3v3KjoD3aj8Y1mkw7DcDhCm+N1T4vFBa76r+9+5FNjpMS0QOkZxx9bDYq3xuqfF4opudJugxA6ZOeOWaB0O4h3f1RDuId39VEv3N7T4Il+5vafBBNwyCSMNw6iN/WmJN9wIvXQD1ndO7qTL43jtQWQq3xvHagPG8dqCyEIQQ7JIs1dtxus3ojaNyeVSzdBv2R+SA07eJvaEadvE3tCYhAuzmRI3u/wAimrLyjY2VqbqdRt5p2SRiDIxBG0BcEejzKbYZReQYkCs7YSZxf1x+CD06F5KpyA0h3+2q62cVyJEER/yZYnDrWmxcjDSsJo1mBpkE1yQC0NiW6Qz0BsP4lB3jUAcZIGAzPW5W07OJvaExCBenZxN7Qo0rS4QQc8imqj+k33oLoQhAIQSq3xvHagsq1GXgRvBHai+N47VWrVAaTgYBOe5BN128dn6ouu3js/VVl+5vePlRL9ze8fKgtddvHZ+qRTYY1mvJxxDuv7SbL9ze8fKksOGsHzjMXoz2IL3PVf3v3Iueq/vfuVZG6p8XiiRuqfF4oL2RsAyy5LjtknIBxO8gfghRYwIdDS2XE62Zy1vehBR3Td/yDLEYtOGzOI25KfvVO6fKoc4X3CKgyN4SWnDZE5RjkrSN9Tsd4II+9U7p8qGF0m7LhA6cjHHLVxUyN9Tsd4Ipl0m7MQOnIxxywQXvP4W94+VF5/C3vHyq2vub2nwRr7m9p8EFCxziLzWwDOZOwjItG9X0LeFvYEXnAgEDHdO4n/xMQL0LeFvYEu00m3cGjNuwcQWhJtglkHe3/IIHISea0+BvdCOa0+BvdCBpVLN0G/ZH5KvNafA3uhOCAQluqQYuk+yPFUdaYza7s/VA9fJfTHltzOVSCCaQLaRJFVzQ7Rh7gGtMEgEuiJkZr6f8oM6/571gtVCzPJcaFJziZJfTaSffvj8lMThExl805Q5UovZUY2tdMhkijVDr9TWpsEmWOJmN4OGxe4/00bdsz2h73gVSL1QODiS1hMh2IGIhaxYaG2hZ42/Mt8c4XTs9elTaG02BjcwGtAGO2Bhioz0wnEZy3oWUW9pyk+5NFb1Xfh4oGqjuk33qzHTsI9qh9MO6QB9olBZCVzVnA3uhHNWcDe6EFbYJaJ4m/wCQV9A3hb2BJr0GC6Q1oN5uIA3hakC9A3hb2BVq2ZpaQGtxBGQ2pyrVfdaTuBPYgpefwt7x8qLz+FvePlU63q/ijW9XtKCLz+FvePlSKZMazng4yADGezVWjW9XtKQxxjWL5xmAYz2YIJ+9U7p8qPvVO6fKpvDfU7HeCLw31Ox3ggmx5OweNY9PM5CQNgw6vYhTYzIJuubrHp5mIF4dRhCBbqgFRwmoMj0SWHD6Jg7sVOlHG/u/tWpCDLpRxv7v7VNOo4k3dYQOlIxxy1cVpQgVefwt7x8qLz+FvePlTUIEOplxF4NgGcydhG4b1bmzOEJqECubM4QgWdnCE1CAQhCAQhCAQhCDDXsZzD3+yR+ZWAh4wLz29uK7pCSbIw43Qg5F5/H8SdZqb3HF5I6iPy2rpc1ZwhWp0Wt6IhBShZ7v0ifaf/E5CEAhCEAhCEFXsBEESFTmzOEJqECubM4QqVbI0tIDRiCO0LQhAq8/hb3j5UXn8Le8fKmoQKvP4W94+VIZWw1nOBxkBpjPZqrYhBl0o4390+VGlHG/unyrUhBnsTgQ4i/0j0wQcIGAP0ULQhB//9k=" height=900 width=900></img></center>

# Overview 

## Setup
We load a few libraries for data visualization here.
```{r}
# general visualisation
library('ggplot2') # visualisation
library('scales') # visualisation
library('patchwork') # visualisation
library('RColorBrewer') # visualisation
library('corrplot') # visualisation

# general data manipulation

library('dplyr') # data manipulation
library('readr') # input/output
library('vroom') # input/output
library('skimr') # overview
library('tibble') # data wrangling
library('tidyr') # data wrangling
library('stringr') # string manipulation
library('forcats') # factor manipulation
```
## Loading our data

```{r echo=FALSE}
if (dir.exists("/kaggle")){
  path <- "/kaggle/input/google-cloud-ncaa-march-madness-2020-division-1-mens-tournament/"
} else {
  path <- ""
}

subpath <- "MDataFiles_Stage1/"
```
```{r}
teams <- vroom(str_c(path, subpath, "MTeams.csv"), col_types = cols())
seasons <- vroom(str_c(path, subpath, "MSeasons.csv"), col_types = cols())
seeds <- vroom(str_c(path, subpath, "MNCAATourneySeeds.csv"), col_types = cols())
regular_res <- vroom(str_c(path, subpath, "MRegularSeasonCompactResults.csv"), col_types = cols())
bracket_res <- vroom(str_c(path, subpath, "MNCAATourneyCompactResults.csv"), col_types = cols())

events19 <- vroom(str_c(path, "MEvents2019.csv"), col_types = cols())

players <- vroom(str_c(path, "MPlayers.csv"), col_types = cols())
```

## Previous events
```{r}
ev15 <- vroom(str_c(path, "MEvents2015.csv"), col_types = cols())
ev16 <- vroom(str_c(path, "MEvents2016.csv"), col_types = cols())
ev17 <- vroom(str_c(path, "MEvents2017.csv"), col_types = cols())
ev18 <- vroom(str_c(path, "MEvents2018.csv"), col_types = cols())
```

# Summaries


## General items {.tabset .tabset-fade .tabset-pills}

### Teams
```{r}
summary(teams)
```

### Seasons
```{r}
summary(seasons)
```

### Players
```{r}
summary(players)
```

## Events over the years {.tabset .tabset-fade .tabset-pills}

### A brief helper
```{r}
library('kableExtra') # will be useful
```

### 2015
```{r}
summary(ev15)
```
### 2016
```{r}
summary(ev16)
```
### 2017
```{r}
summary(ev17)
```
### 2018
```{r}
summary(ev18)
```
### 2019
```{r}
summary(events19)
```

## Data Samples 
### Using kable {.tabset .tabset-fade}
#### Teams

```{r}
teams %>% 
  head(5) %>% 
  kable() %>% 
  kable_styling()
```


```{r}
teams %>% 
  tail(5) %>% 
  kable() %>% 
  kable_styling()
```

#### Seasons

```{r}
seasons %>% 
  head(5) %>% 
  kable() %>% 
  kable_styling()
```

```{r}
seasons %>% 
  tail(5) %>% 
  kable() %>% 
  kable_styling()
```

#### Players

```{r}
players %>% 
  head(5) %>% 
  kable() %>% 
  kable_styling()
```

```{r}
players %>% 
  tail(5) %>% 
  kable() %>% 
  kable_styling()
```

#### Events for 2015

```{r}
ev15 %>% 
  head(5) %>% 
  kable() %>% 
  kable_styling()
```


```{r}
ev15 %>% 
  tail(5) %>% 
  kable() %>% 
  kable_styling()
```

#### Events for 2016

```{r}
ev16 %>% 
  head(5) %>% 
  kable() %>% 
  kable_styling()
```

```{r}
seasons %>% 
  tail(5) %>% 
  kable() %>% 
  kable_styling()
```

#### Events for 2017

```{r}
ev17 %>% 
  head(5) %>% 
  kable() %>% 
  kable_styling()
```

```{r}
ev17 %>% 
  tail(5) %>% 
  kable() %>% 
  kable_styling()
```

#### Events for 2018
```{r}
ev18 %>% 
  head(5) %>% 
  kable() %>% 
  kable_styling()
```

```{r}
ev18 %>% 
  tail(5) %>% 
  kable() %>% 
  kable_styling()
```

#### Events for 2019
```{r}
events19 %>% 
  head(5) %>% 
  kable() %>% 
  kable_styling()
```

```{r}
events19 %>% 
  tail(5) %>% 
  kable() %>% 
  kable_styling()
```

### Where is the missing data?

We should first see what exactly is missing before proceeding to (evil laughs) **devastate it utterly!**

```{r}
library(tidyverse)

library(ggplot2)

library(caret)
install.packages("caretEnsemble")
library(caretEnsemble)

library(psych)

library(Amelia)

library(mice)

library(GGally)

library(rpart)

library(randomForest)

```{r}
missmap(teams)
```

```{r}
missmap(seasons)
```


```{r}
missmap(players)
```

```{r}
missmap(regular_res)
```

```{r}
missmap(bracket_res)
```

```{r}
missmap(ev15)
```


```{r}
missmap(ev16)
```

```{r}
missmap(ev17)
```

```{r}
missmap(ev18)
```

```[r]
missmap(ev19)
```

## Eliminating missing data

We are going to use the package `mice` to eliminate the missing data.

First, we must load the library:
```{r}
library(mice)
```

Now, we must remove missing data:
```{r}
mice_teams <- mice(teams[, c("TeamID","TeamName","FirstD1Season","LastD1Season")], method='rf')
mice_complete1 <- complete(mice_teams)
```

```{r}
teams$TeamID <- mice_complete1$TeamID
teams$TeamName <- mice_complete1$TeamName
teams$FirstD1Season <- mice_complete1$FirstD1Season
teams$LastD1Season<- mice_complete1$LastD1Season
```

```{r}
mice_S <- mice(seasons[, c("DayZero","Season", "RegionW","RegionX", "RegionY", "RegionZ")], method='rf')
mice_complete1 <- complete(mice_S)
```

```{r}
seasons$DayZero <- mice_S$DayZero
seasons$Season <- mice_S$Season
seasons$RegionW <- mice_S$RegionW
seasons$RegionX <- mice_S$RegionX
seasons$RegionY <- mice_S$RegionY
seasons$RegionZ <- mice_S$RegionZ
```

```{r}
PL <- mice(players[, c("PlayerID","LastName","FirstName", "TeamID")], method='rf')
mice_complete1 <- complete(PL)
```



## Relationships with treemaps {.tabset .tabset-fade .tabset-pills}
### Helper
```{r}
library(treemap)

treemapper = function(data, x, y) {
    tm <- treemap(data, index = x,
              vSize = y)
}
```
### Seasons {.tabset .tabset-fade .tabset-pills}
#### Regular Season Results {.tabset .tabset-fade .tabset-pills}
##### Winner Score and Loser score
```{r}
value_data = regular_res %>% 
      select(WScore) %>%
      group_by(WScore) %>% 
      summarise(LScore = n())%>%
                filter(LScore < 4)

```
```{r}
treemapper(value_data, 'WScore', 'LScore')
```
```{r}
value_data = regular_res %>% 
      select(LScore) %>%
      group_by(LScore) %>% 
      summarise(WScore = n())%>%
                filter(WScore < 4)

```
```{r}
treemapper(value_data, 'LScore', 'WScore')
```

#### Bracketed Results {.tabset .tabset-fade .tabset-pills}
##### Winner Score and Loser score

```{r}
value_data = bracket_res %>% 
      select(WScore) %>%
      group_by(WScore) %>% 
      summarise(LScore = n())%>%
                filter(LScore < 4)

```
```{r}
treemapper(value_data, 'WScore', 'LScore')
```

##### Location
##### Winners and Loser

```{r}
value_data = bracket_res %>% 
      select(WTeamID) %>%
      group_by(WTeamID) %>% 
      summarise(LTeamID  = n())%>%
                filter(LTeamID  < 4)

```
```{r}
treemapper(value_data, 'WTeamID', 'LTeamID')
```

```{r}
value_data = bracket_res %>% 
      select(LTeamID) %>%
      group_by(LTeamID) %>% 
      summarise(WTeamID  = n())%>%
                filter(WTeamID  < 4)

```
```{r}
treemapper(value_data, 'LTeamID', 'WTeamID')
```

# Correlations with corrgram


In this part of my data analysis, I will use the <code>correlogram</code> package to create correlograms.

## Loading corrgram
```{r}
library(corrgram)
```

## Analysis 1: Bracketed and Regular Results

### Bracketed results


```{r}
corrgram(bracket_res, order=TRUE, lower.panel=panel.shade,
  upper.panel=panel.pie, text.panel=panel.txt,
  main="Bracketed Results")
```

#### Another analysis
```{r}
library(corrgram)
col.corrgram <- function(ncol){   
  colorRampPalette(c("darkgoldenrod4", "burlywood1",
  "darkkhaki", "darkgreen"))(ncol)}
corrgram(bracket_res, order=TRUE, lower.panel=panel.shade,
   upper.panel=panel.pie, text.panel=panel.txt,
   main="More")
```

### Regular results

```{r}
corrgram(regular_res, order=TRUE, lower.panel=panel.shade,
  upper.panel=panel.pie, text.panel=panel.txt,
  main="Regular Results")
```

#### Another analysis
```{r}
library(corrgram)
col.corrgram <- function(ncol){   
  colorRampPalette(c("darkgoldenrod4", "burlywood1",
  "darkkhaki", "darkgreen"))(ncol)}
corrgram(regular_res, order=TRUE, lower.panel=panel.shade,
   upper.panel=panel.pie, text.panel=panel.txt,
   main="another one")
```

---

Out of everything, we can see two main things over here:
<ol>
<li>In both bracketed and regular season results, we can view that <strong>there is an extremely strong correlation between the winner and the loser's score</strong>.
<li>We can also see that in regular results, there is absolutely no correlation between LTeamID and NumScore.
</ol>

### Takeaways: Correlations between bracketed and regular results 
#### Top 3 correlations in regular seasons
<ol>
<li>WScore + LScore
<li>Season + DayNum
<li>Season + LScore
</ol>

#### Top 3 correlations in bracketed seasons
<ol>
<li>WScore + LScore
<li>Season + DayNum
<li>Season + LScore
</ol>

## Analysis 2: Teams

```{r}
corrgram(teams, order=TRUE, lower.panel=panel.shade,
  upper.panel=panel.pie, text.panel=panel.txt,
  main="Teams")
```

## Takeaways 2: Teams

### Top 3 correlations in teams
<ol>
<li>TeamID + LastD1Season
<li>TeamID + FirstD1Season
<li>FirstD1Season + LastD1Season
</ol>


## Analysis 3: Seeds

```{r}
corrgram(seeds, order=TRUE, lower.panel=panel.shade,
  upper.panel=panel.pie, text.panel=panel.txt,
  main="Bracketed Results")
```

---

## Takeaways 3: Seasons

### Top correlation
<ol><li>Season + TeamID</ol>

# Individual Feature Analysis

We're going to use a regression analysis here.

## First analysis: Teams {.tabset .tabset-fade }

### Preparations
```{r}
install.packages("ggeffects")
install.packages("ggalluvial")
library(ggalluvial)
library(ggeffects)
library(viridis)
library(knitr)
library(kableExtra)
library(countrycode)
library(highcharter)
library(tidyverse)
library(magrittr)
```

### Starting off
```{r}
teams %>%
  select(FirstD1Season, LastD1Season, TeamName, TeamID) %>%
  drop_na() %>%
  filter(FirstD1Season == "1985" | FirstD1Season == "2020") %>%
  group_by(FirstD1Season, LastD1Season, TeamName, TeamID) %>%
  summarize(freq = n()) %>%
  ungroup() %>% 
  ggplot(aes(y = freq, axis1 = FirstD1Season, axis2 = LastD1Season, axis3 = TeamName, axis4 = TeamID)) + 
  geom_alluvium(aes(fill = FirstD1Season), width = 1/12) + 
  geom_stratum(width = 1/12, fill = "black", color = "grey") + 
  geom_label(stat = "stratum", label.strata = TRUE) +
  scale_x_discrete(limits = c("First Season", "Last Season", "TeamName", "Team ID"), expand = c(.11, .01)) + 
  labs(x = "", y = "") + 
  theme_minimal() + 
  theme(legend.position="none") 
```

So we can see **a positive bunch** of uniques in the TeamName column. This is something that was expected. 

Let's try a switcharound to see what we get with the other variables, maybe it will be possible to have some fun and/or have some bets for you bookies in there.


```{r}
teams %>%
  select(FirstD1Season, LastD1Season, TeamID) %>%
  drop_na() %>%
  filter(LastD1Season == "1985" | LastD1Season == "2020") %>%
  group_by(LastD1Season, FirstD1Season, TeamID) %>%
  summarize(freq = n()) %>%
  ungroup() %>% 
  ggplot(aes(y = freq, axis1 = LastD1Season, axis2 = FirstD1Season, axis4 = TeamID)) + 
  geom_alluvium(aes(fill = FirstD1Season), width = 1/12) + 
  geom_stratum(width = 1/12, fill = "black", color = "grey") + 
  geom_label(stat = "stratum", label.strata = TRUE) +
  scale_x_discrete(limits = c("First Season", "Last Season", "Team ID"), expand = c(.11, .01)) + 
  labs(x = "", y = "") + 
  theme_minimal() + 
  theme(legend.position="none") 
```

Also, now we get a chance to look at regular results and bracket results. (briefly). A more indepth explanation is coming soon.

First, we'll look at the winning teams.

#### Winners
```{r}
regular_res %>% 
  ggplot(aes(x=WTeamID)) +
  geom_histogram(alpha = 0.5, fill = "#5EB296", colour = "#4D4D4D") +
  scale_x_continuous(labels = comma) +
  scale_y_continuous(labels = comma) +
  ggtitle("Winning Team Ids ", subtitle = "We are going to look at legendary winners through the years") +
  labs(x= "Winning Team", y= "Count")
```

---
#### Insights from winning teams:

* One team has a close lead when it comes to wins. (approximately 6,600 wins)
* Another team has a niche when it comes to the least wins, as it has only 4,200.
* The range of the data is:
$range = max - min
$ = 6600 - 4200
$ = 4400

#### Losers

Now, a brief look at the legendary losers.
```{r}
regular_res %>% 
  ggplot(aes(x=LTeamID)) +
  geom_histogram(alpha = 0.5, fill = "#5EB296", colour = "#4D4D4D") +
  scale_x_continuous(labels = comma) +
  scale_y_continuous(labels = comma) +
  ggtitle("Losing Team Ids ", subtitle = "We are going to look at legendary losers through the years") +
  labs(x= "Winning Team", y= "Count")
```


#### Insights from losing teams:

### Target 1: Bracketed Results

```{r}
library(tidyverse)
library(gridExtra)
library(knitr)
library(magrittr)
library(ggExtra)
```

```{r}
# Join team names to tourney compact dataset
tourney_stats_compact <- bracket_res %>%
  left_join(teams, by = c("WTeamID" = "TeamID")) %>%
  left_join(teams, by = c("LTeamID" = "TeamID"))

tourney_stats_compact <- tourney_stats_compact %>%
  rename(WTeamName = TeamName.x,
         LTeamName = TeamName.y)

tourney_stats_compact$season_day <- paste(tourney_stats_compact$Season, tourney_stats_compact$DayNum, sep = "_")


# then create a feature to label the round of the tournament
tourney_stats_compact <- tourney_stats_compact %>%
  mutate(TourneyRound = ifelse(DayNum %in% c(136, 137), "First Round", ifelse(DayNum %in% c(138, 139), "Second Round", ifelse(DayNum %in% c(143, 144), "Sweet 16", ifelse(DayNum %in% c(145, 146), "Elite 8", ifelse(DayNum == 152, "Final Four", "Championship Game")))))) %>%
  mutate(TourneyRound = factor(TourneyRound, levels = c("First Round", "Second Round", "Sweet 16", "Elite 8", "Final Four", "Championship Game")))

ncaa_champs <- tourney_stats_compact %>%
  group_by(Season) %>%
  summarise(max_days = max(DayNum)) %>%
  mutate(season_day = paste(Season, max_days, sep = "_")) %>%
  left_join(tourney_stats_compact, by = "season_day") %>% ungroup() %>%
  select(-Season.y) %>%
  rename(Season = Season.x)

win_plot <- ncaa_champs %>%
  group_by(WTeamName) %>%
  summarise(n = n()) %>%
  ggplot(aes(x=reorder(WTeamName,n), y=n)) +
  geom_bar(stat = "identity", fill = "steelblue", color = "black") +
  theme_minimal()+
  labs(title = "Most Tourney Wins since 1985", x= "Winner", y= "Titles") +
  coord_flip()


lose_plot <- ncaa_champs %>%
  group_by(LTeamName) %>%
  summarise(n = n()) %>%
  ggplot(aes(x=reorder(LTeamName, n), y=n)) +
  geom_bar(stat = "identity", fill = "orange", colour = "black") +
  theme_minimal()+
  labs(title = "Most Tourney Bridesmaids since 1985", x= "Runner-up", y= "Titles") +
  coord_flip()

grid.arrange(win_plot, lose_plot, ncol = 2)
```
---

Now we're going to work with the main focus of ths kernel: **analyzing the bracketed and regular season results**. My aim is to create a prime analysis for the bracketed and regular results.

# The Primary Analysis: Part 1

In the primary analysis, we will mainly look at the results in the regular and bracketed seasons.

## Part 1: Regular seasons

So, in this part of our analysis, we shall look at the regular season results of years long gone.

### An introduction

The tournament consists of several rounds. They are currently named, in order of first to last:

* The First Four
* The First Round (the Round of 64)
* The Second Round (the Round of 32)
* The Regional Semi-finals (participating teams are known popularly as the "Sweet Sixteen")
* The Regional Finals (participating teams are known commonly as the "Elite Eight")
* The National Semi-finals (participating teams are referred to officially as the "Final Four")
* The National Championship

The tournament is single-elimination, which increases the chance of an underdog and lower-seeded "Cinderella team" advancing to subsequent rounds. Although these lower-ranked teams are forced to play stronger teams, they need only one win to advance (instead of needing to win a majority of games in a series, as in professional basketball).


The University of Dayton Arena, which has hosted all First Four games since the round's inception in 2011, as well as its precursor, the single "play-in" game held from 2001 to 2010. As of 2019, the arena has hosted 123 tournament games, the most of any venue.
First held during 2011, the First Four are games between the four lowest-ranked at-large teams and the four lowest-ranked automatic-bid (conference-champion) teams.

### A brief reminder or two

#### Events for 2015

```{r}
ev15 %>% 
  head(5) %>% 
  kable() %>% 
  kable_styling()
```


```{r}
ev15 %>% 
  tail(5) %>% 
  kable() %>% 
  kable_styling()
```

#### Events for 2016

```{r}
ev16 %>% 
  head(5) %>% 
  kable() %>% 
  kable_styling()
```

```{r}
seasons %>% 
  tail(5) %>% 
  kable() %>% 
  kable_styling()
```

#### Events for 2017

```{r}
ev17 %>% 
  head(5) %>% 
  kable() %>% 
  kable_styling()
```

```{r}
ev17 %>% 
  tail(5) %>% 
  kable() %>% 
  kable_styling()
```

#### Events for 2018
```{r}
ev18 %>% 
  head(5) %>% 
  kable() %>% 
  kable_styling()
```

```{r}
ev18 %>% 
  tail(5) %>% 
  kable() %>% 
  kable_styling()
```

#### Events for 2019
```{r}
events19 %>% 
  head(5) %>% 
  kable() %>% 
  kable_styling()
```

```{r}
events19 %>% 
  tail(5) %>% 
  kable() %>% 
  kable_styling()
```

### A pairplot to look at each season {.tabset .tabset-fade .tabset-pills}

We're going to use the function `ggpairs` from ggplot2 to look at the data concerning each season.

#### 2015
```{r}
ggpairs(head(ev15))
```

#### 2016
```{r}
ggpairs(head(ev16))
```

#### 2017
```{r}
ggpairs(head(ev17))
```

#### 2018
```{r}
ggpairs(head(ev18))
```

#### 2019
```{r}
ggpairs(head(events19))
```



### What scores do we have?

```{r}
regular_res %>% 
  ggplot(aes(x=LScore)) +
  geom_histogram(alpha = 0.5, fill = "#5EB296", colour = "#4D4D4D") +
  scale_x_continuous(labels = comma) +
  scale_y_continuous(labels = comma) +
  ggtitle("Loser's Score", subtitle = "What did the losers do?") +
  labs(x= "Loser Score", y= "Count")
```

```{r}
regular_res %>% 
  ggplot(aes(x=WScore)) +
  geom_histogram(alpha = 0.5, fill = "#5EB296", colour = "#4D4D4D") +
  scale_x_continuous(labels = comma) +
  scale_y_continuous(labels = comma) +
  ggtitle("Winning Score ", subtitle = "Winners are doing better") +
  labs(x= "Winning Score", y= "Count")
```

#### Observations

We see that:

* A lot of victories have a score of about 75 to 80. 
* A lot of losses have a score of about 50.

### Using corrplot {.tabset .tabset-fade}

We will use the library `corrplot` to visualize correlations between the data for:
* Teams
* Seasons
* (main) Bracketed and regular results
* Seeds
* Events over the years

#### Import

```{r}
library(corrplot)
```

#### Bracketed results {.tabset .tabset-fade}

We are going to use different variables to test the correlation of bracketed results: specifically <i>WScore and LScore</i>. 

##### WScore-based corrplot

We are going to use WScore as the base for this correlation plot.

```{r}
numericVars <- which(sapply(bracket_res, is.numeric)) #index vector numeric variables
numericVarNames <- names(numericVars) #saving names vector for use later on
cat('There are', length(numericVars), 'numeric variables')
## There are 37 numeric variables
all_numVar <- bracket_res[, numericVars]
cor_numVar <- cor(all_numVar, use="pairwise.complete.obs") #correlations of all numeric variables

#sort on decreasing correlations with SalePrice
cor_sorted <- as.matrix(sort(cor_numVar[,'WScore'], decreasing = TRUE))
 #select only high corelations
CorHigh <- names(which(apply(cor_sorted, 1, function(x) abs(x)>0.5)))
cor_numVar <- cor_numVar[CorHigh, CorHigh]

corrplot.mixed(cor_numVar, tl.col="black", tl.pos = "lt")
```

##### LScore-based corrplot

We are going to use LScore as the base for this correlation plot.                             
                             
```{r}
numericVars <- which(sapply(bracket_res, is.numeric)) #index vector numeric variables
numericVarNames <- names(numericVars) #saving names vector for use later on
cat('There are', length(numericVars), 'numeric variables')

all_numVar <- bracket_res[, numericVars]
cor_numVar <- cor(all_numVar, use="pairwise.complete.obs") #correlations of all numeric variables

#sort on decreasing correlations with SalePrice
cor_sorted <- as.matrix(sort(cor_numVar[,'LScore'], decreasing = TRUE))
 #select only high corelations
CorHigh <- names(which(apply(cor_sorted, 1, function(x) abs(x)>0.5)))
cor_numVar <- cor_numVar[CorHigh, CorHigh]

corrplot.mixed(cor_numVar, tl.col="black", tl.pos = "lt")
```    
                             
---
            

#### Takeaways
                             
1. **WScore and LScore both have exactly the same correlation.**
          
                                                        
                            
#### Regular results {.tabset .tabset-fade}

We are going to use different variables to test the correlation of regular results: specifically <i>WScore and LScore</i>. 

##### WScore-based corrplot

We are going to use WScore as the base for this correlation plot.

```{r}
numericVars <- which(sapply(regular_res, is.numeric)) #index vector numeric variables
numericVarNames <- names(numericVars) #saving names vector for use later on
cat('There are', length(numericVars), 'numeric variables')
## There are 37 numeric variables
all_numVar <- bracket_res[, numericVars]
cor_numVar <- cor(all_numVar, use="pairwise.complete.obs") #correlations of all numeric variables

#sort on decreasing correlations with SalePrice
cor_sorted <- as.matrix(sort(cor_numVar[,'WScore'], decreasing = TRUE))
 #select only high corelations
CorHigh <- names(which(apply(cor_sorted, 1, function(x) abs(x)>0.5)))
cor_numVar <- cor_numVar[CorHigh, CorHigh]

corrplot.mixed(cor_numVar, tl.col="black", tl.pos = "lt")
```

##### LScore-based corrplot

We are going to use LScore as the base for this correlation plot.                             
                             
```{r}
numericVars <- which(sapply(regular_res, is.numeric)) #index vector numeric variables
numericVarNames <- names(numericVars) #saving names vector for use later on
cat('There are', length(numericVars), 'numeric variables')

all_numVar <- bracket_res[, numericVars]
cor_numVar <- cor(all_numVar, use="pairwise.complete.obs") #correlations of all numeric variables

#sort on decreasing correlations with SalePrice
cor_sorted <- as.matrix(sort(cor_numVar[,'LScore'], decreasing = TRUE))
 #select only high corelations
CorHigh <- names(which(apply(cor_sorted, 1, function(x) abs(x)>0.5)))
cor_numVar <- cor_numVar[CorHigh, CorHigh]

corrplot.mixed(cor_numVar, tl.col="black", tl.pos = "lt")
```                        
                             
                             
---
            

#### Takeaways
   
* Same as the previous one.
                             
## Event correlation for previous years {.tabset .tabset-fade}
                             
We are going to visualize the event correlation for previous years.
                             
We are going to use different variables to test the correlation of bracketed results for events.
                             
        
                             
---
            

#### Takeaways                             
                    
                             
**Note: The following section will not just be used for the Analytics competitions, but also for the NCAAM competition. I will be using LightGBM to predict and make observations from lightGBM predictions.**

# Solutions

 
This part contains my solution for the NCAAM competition. 
<br>
It uses:
    * LightGBM
    * XGBoost
stacked together.

## Credits for the solution
        * Larxel: https://www.kaggle.com/andrewmvd/lightgbm-in-r
        * Andrew Lukyanenko: https://www.kaggle.com/artgor/march-madness-2020-ncaam-eda-and-baseline

## Imports

```{r}
library(lightgbm)
library(xgboost)
library(MLmetrics)
library(data.table)
library(Matrix)
```

```{r}
set.seed(257)
```



Working on it.