{
  "id": 358786,
  "title": "Privileged ball positions",
  "url": "/competitions/tabular-playground-series-oct-2022/discussion/358786",
  "author_name": "",
  "post_date": "2022-10-09T13:50:43.793363Z",
  "votes": 19,
  "comment_count": 16,
  "views": 0,
  "content": "<p>As many of you will have deduced, the closer the ball is to the opponent's goal, the greater the probability of scoring a goal. </p>\n<p>A logical question that may arise is: <br>\n<strong>What does this probability distribution look like if we only look at the position of the ball and completely ignore the location of the player?</strong> </p>\n<p>Well, if we analyse all the training data we get a distribution for each team that looks something like this:</p>\n<p>Team A:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4696479%2Fff0b3602d1033659c2ce68ae341b3865%2F__results___3_2.png?generation=1665322117461398&amp;alt=media\" alt=\"\"></p>\n<p>Team B:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4696479%2F0bcc2e04ac384e124bb5f533585b3eec%2F__results___3_4.png?generation=1665322163197539&amp;alt=media\" alt=\"\"></p>\n<p>As we can see, apart from the goal area there are two privileged positions of high goal probability: the kick-off zone and the touchline at the back of the pitch.</p>\n<p>I have published these two dictionaries so that you can use this information as an additional feature:<br>\n<a href=\"https://www.kaggle.com/datasets/maherelouahabi/court-probability-distrubution/\" target=\"_blank\">https://www.kaggle.com/datasets/maherelouahabi/court-probability-distrubution/</a></p>\n<p>How to use it:</p>\n<pre><code>import pickle\nimport pandas as pd\n\ndf = pd.read_csv('...')\n\ndef open_pkl(name):\n    with open(f'{name}.pkl', 'rb') as f:\n        return pickle.load(f)\n\ndict_A = open_pkl('../input/court-probability/dict_A')\n\ndict_B = open_pkl('../input/court-probability/dict_B')\n\ndf['ball_pos_prob_A'] = df.apply(lambda x: dict_A.get(f'{(round(x.ball_pos_y),round(x.ball_pos_x))}',0), axis=1)\ndf['ball_pos_prob_B'] = df.apply(lambda x: dict_B.get(f'{(round(x.ball_pos_y),round(x.ball_pos_x))}',0), axis=1)\n</code></pre>\n<p><strong>(UPDATE)</strong><br>\nI have corrected the distribution calculation and updated both the dictionaries and the plots (Thank you <a href=\"https://www.kaggle.com/siukeitin\" target=\"_blank\">@siukeitin</a>). I leave the previous plots so that anyone can learn from my mistakes. </p>\n<p>Team A:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4696479%2F89fefdbc1a68ef0c2dd824ad1c59c3f0%2FA.png?generation=1665569798239607&amp;alt=media\" alt=\"\"></p>\n<p>Team B:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4696479%2F32625d5a599b2208e2c8dbada98c0ba6%2FB.png?generation=1665569813790580&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "1979475",
      "postDate": "10/09/2022 13:50:43",
      "content": "<p>As many of you will have deduced, the closer the ball is to the opponent's goal, the greater the probability of scoring a goal. </p>\n<p>A logical question that may arise is: <br>\n<strong>What does this probability distribution look like if we only look at the position of the ball and completely ignore the location of the player?</strong> </p>\n<p>Well, if we analyse all the training data we get a distribution for each team that looks something like this:</p>\n<p>Team A:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4696479%2Fff0b3602d1033659c2ce68ae341b3865%2F__results___3_2.png?generation=1665322117461398&amp;alt=media\" alt=\"\"></p>\n<p>Team B:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4696479%2F0bcc2e04ac384e124bb5f533585b3eec%2F__results___3_4.png?generation=1665322163197539&amp;alt=media\" alt=\"\"></p>\n<p>As we can see, apart from the goal area there are two privileged positions of high goal probability: the kick-off zone and the touchline at the back of the pitch.</p>\n<p>I have published these two dictionaries so that you can use this information as an additional feature:<br>\n<a href=\"https://www.kaggle.com/datasets/maherelouahabi/court-probability-distrubution/\" target=\"_blank\">https://www.kaggle.com/datasets/maherelouahabi/court-probability-distrubution/</a></p>\n<p>How to use it:</p>\n<pre><code>import pickle\nimport pandas as pd\n\ndf = pd.read_csv('...')\n\ndef open_pkl(name):\n    with open(f'{name}.pkl', 'rb') as f:\n        return pickle.load(f)\n\ndict_A = open_pkl('../input/court-probability/dict_A')\n\ndict_B = open_pkl('../input/court-probability/dict_B')\n\ndf['ball_pos_prob_A'] = df.apply(lambda x: dict_A.get(f'{(round(x.ball_pos_y),round(x.ball_pos_x))}',0), axis=1)\ndf['ball_pos_prob_B'] = df.apply(lambda x: dict_B.get(f'{(round(x.ball_pos_y),round(x.ball_pos_x))}',0), axis=1)\n</code></pre>\n<p><strong>(UPDATE)</strong><br>\nI have corrected the distribution calculation and updated both the dictionaries and the plots (Thank you <a href=\"https://www.kaggle.com/siukeitin\" target=\"_blank\">@siukeitin</a>). I leave the previous plots so that anyone can learn from my mistakes. </p>\n<p>Team A:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4696479%2F89fefdbc1a68ef0c2dd824ad1c59c3f0%2FA.png?generation=1665569798239607&amp;alt=media\" alt=\"\"></p>\n<p>Team B:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4696479%2F32625d5a599b2208e2c8dbada98c0ba6%2FB.png?generation=1665569813790580&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "As many of you will have deduced, the closer the ball is to the opponent's goal, the greater the probability of scoring a goal. \n\nA logical question that may arise is: \n**What does this probability distribution look like if we only look at the position of the ball and completely ignore the location of the player?** \n\nWell, if we analyse all the training data we get a distribution for each team that looks something like this:\n\nTeam A:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4696479%2Fff0b3602d1033659c2ce68ae341b3865%2F__results___3_2.png?generation=1665322117461398&alt=media)\n\nTeam B:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4696479%2F0bcc2e04ac384e124bb5f533585b3eec%2F__results___3_4.png?generation=1665322163197539&alt=media)\n\nAs we can see, apart from the goal area there are two privileged positions of high goal probability: the kick-off zone and the touchline at the back of the pitch.\n\nI have published these two dictionaries so that you can use this information as an additional feature:\nhttps://www.kaggle.com/datasets/maherelouahabi/court-probability-distrubution/\n\nHow to use it:\n```\nimport pickle\nimport pandas as pd\n\ndf = pd.read_csv('...')\n\ndef open_pkl(name):\n    with open(f'{name}.pkl', 'rb') as f:\n        return pickle.load(f)\n\ndict_A = open_pkl('../input/court-probability/dict_A')\n\ndict_B = open_pkl('../input/court-probability/dict_B')\n\ndf['ball_pos_prob_A'] = df.apply(lambda x: dict_A.get(f'{(round(x.ball_pos_y),round(x.ball_pos_x))}',0), axis=1)\ndf['ball_pos_prob_B'] = df.apply(lambda x: dict_B.get(f'{(round(x.ball_pos_y),round(x.ball_pos_x))}',0), axis=1)\n```\n\n**(UPDATE)**\nI have corrected the distribution calculation and updated both the dictionaries and the plots (Thank you @siukeitin). I leave the previous plots so that anyone can learn from my mistakes. \n\nTeam A:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4696479%2F89fefdbc1a68ef0c2dd824ad1c59c3f0%2FA.png?generation=1665569798239607&alt=media)\n\nTeam B:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4696479%2F32625d5a599b2208e2c8dbada98c0ba6%2FB.png?generation=1665569813790580&alt=media)",
      "votes": null
    },
    {
      "id": "1979551",
      "postDate": "10/09/2022 14:43:17",
      "content": "<p>Really an insightful post, I will definitely try to the dictionaries as features, I may even try and remove the real ball position from the features used in a model. I have the feeling that many models are not working to their full potential since the position of players and positions of balls are hard to use as they are.</p>\n<p>A really important aspect to notice is that the distance is not the only aspect that makes up the likelihood, also being or not at an edge of the field seems to contribute, another insight that may be used for FE. </p>\n<p>Thank you <a href=\"https://www.kaggle.com/maherelouahabi\" target=\"_blank\">@maherelouahabi</a> ! Keep kaggling and let's enjoy this month's TPS</p>",
      "rawMarkdown": "Really an insightful post, I will definitely try to the dictionaries as features, I may even try and remove the real ball position from the features used in a model. I have the feeling that many models are not working to their full potential since the position of players and positions of balls are hard to use as they are.\n\nA really important aspect to notice is that the distance is not the only aspect that makes up the likelihood, also being or not at an edge of the field seems to contribute, another insight that may be used for FE. \n\nThank you @maherelouahabi ! Keep kaggling and let's enjoy this month's TPS",
      "votes": null
    },
    {
      "id": "1979585",
      "postDate": "10/09/2022 15:03:05",
      "content": "<p>Great observations <a href=\"https://www.kaggle.com/maherelouahabi\" target=\"_blank\">@maherelouahabi</a>, thanks for sharing this information too!</p>",
      "rawMarkdown": "Great observations @maherelouahabi, thanks for sharing this information too!",
      "votes": null
    },
    {
      "id": "1979628",
      "postDate": "10/09/2022 15:31:26",
      "content": "<p>You are welcome. Happy kaggling <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> </p>",
      "rawMarkdown": "You are welcome. Happy kaggling @ravi20076",
      "votes": null
    },
    {
      "id": "1979634",
      "postDate": "10/09/2022 15:37:40",
      "content": "<p>Dear Pietro,</p>\n<p>I think one of the main problem with the positions is the high factorial cardinality of them. Approaches like this one, cook this high cardinality.<br>\nVery interesting what you say about the other factors that increase likelihood. I will think about it.</p>\n<p>Happy kaggling <a href=\"https://www.kaggle.com/pietromaldini1\" target=\"_blank\">@pietromaldini1</a> <br>\nRegards</p>",
      "rawMarkdown": "Dear Pietro,\n\nI think one of the main problem with the positions is the high factorial cardinality of them. Approaches like this one, cook this high cardinality.\nVery interesting what you say about the other factors that increase likelihood. I will think about it.\n\nHappy kaggling @pietromaldini1 \nRegards",
      "votes": null
    },
    {
      "id": "1981261",
      "postDate": "10/10/2022 17:33:00",
      "content": "<p>Amazing insight! In soccer, there is a thing called Expected Goal that match the probability of scoring with an area of the field, some people even use the position of the ball to optimize this probabilities</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2224301%2Fda46c492db121d5bf0b14b0c17ea10b2%2Ffutbol_prob.jpg?generation=1665423082433728&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Amazing insight! In soccer, there is a thing called Expected Goal that match the probability of scoring with an area of the field, some people even use the position of the ball to optimize this probabilities\n\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2224301%2Fda46c492db121d5bf0b14b0c17ea10b2%2Ffutbol_prob.jpg?generation=1665423082433728&alt=media)",
      "votes": null
    },
    {
      "id": "1981529",
      "postDate": "10/10/2022 20:38:37",
      "content": "<p>It seems that these features are related to target encoding the integer ball coordinates with respect to the two target variables. For example,</p>\n<pre><code>from category_encoders.target_encoder import TargetEncoder\n\ntrain['int_ball_pos']=train.ball_pos_x.astype('int').astype('str')+'_'+train.ball_pos_y.astype('int').astype('str')\nenc_A = TargetEncoder(return_df=False).fit(train[['int_ball_pos']],train.team_A_scoring_within_10sec)\nenc_B = TargetEncoder(return_df=False).fit(train[['int_ball_pos']],train.team_B_scoring_within_10sec)\ntrain['ball_pos_prob_A']=enc_A.transform(train[['int_ball_pos']])\ntrain['ball_pos_prob_B']=enc_B.transform(train[['int_ball_pos']])\ntrain.drop(['int_ball_pos'],axis=1,inplace=True)\n</code></pre>\n<p>Why does ball position \\((0,0)\\) have high probability of scoring within the next 10 sec for either team? </p>",
      "rawMarkdown": "It seems that these features are related to target encoding the integer ball coordinates with respect to the two target variables. For example,\n```\nfrom category_encoders.target_encoder import TargetEncoder\n\ntrain['int_ball_pos']=train.ball_pos_x.astype('int').astype('str')+'_'+train.ball_pos_y.astype('int').astype('str')\nenc_A = TargetEncoder(return_df=False).fit(train[['int_ball_pos']],train.team_A_scoring_within_10sec)\nenc_B = TargetEncoder(return_df=False).fit(train[['int_ball_pos']],train.team_B_scoring_within_10sec)\ntrain['ball_pos_prob_A']=enc_A.transform(train[['int_ball_pos']])\ntrain['ball_pos_prob_B']=enc_B.transform(train[['int_ball_pos']])\ntrain.drop(['int_ball_pos'],axis=1,inplace=True)\n```\nWhy does ball position \\\\((0,0)\\\\) have high probability of scoring within the next 10 sec for either team?",
      "votes": null
    },
    {
      "id": "1981556",
      "postDate": "10/10/2022 21:16:46",
      "content": "<p>Dear Pastor,</p>\n<p>Very interesting what you say. I didn't know that a similar concept exists in soccer. Thank you very much for the contribution.</p>\n<p>Happy kaggling <a href=\"https://www.kaggle.com/pastorsoto\" target=\"_blank\">@pastorsoto</a> <br>\nRegards</p>",
      "rawMarkdown": "Dear Pastor,\n\nVery interesting what you say. I didn't know that a similar concept exists in soccer. Thank you very much for the contribution.\n\nHappy kaggling @pastorsoto \nRegards",
      "votes": null
    },
    {
      "id": "1981558",
      "postDate": "10/10/2022 21:21:39",
      "content": "<p>Dear Brococoli,</p>\n<p>I was waiting for this quesiton. The main reason is because it is the point at which the match starts + the point at which they start after scoring a goal. If you look at any gameplay, as soon as the match starts, the players closest to the position of the ball, go as fast as possible to the ball. The team that arrives first, usually transmits a lot of inertia to the ball in the direction of the opponent's goal, often ending in a goal.</p>\n<p>Happy kaggling <a href=\"https://www.kaggle.com/siukeitin\" target=\"_blank\">@siukeitin</a> <br>\nRegards</p>",
      "rawMarkdown": "Dear Brococoli,\n\nI was waiting for this quesiton. The main reason is because it is the point at which the match starts + the point at which they start after scoring a goal. If you look at any gameplay, as soon as the match starts, the players closest to the position of the ball, go as fast as possible to the ball. The team that arrives first, usually transmits a lot of inertia to the ball in the direction of the opponent's goal, often ending in a goal.\n\nHappy kaggling @siukeitin \nRegards",
      "votes": null
    },
    {
      "id": "1981559",
      "postDate": "10/10/2022 21:22:40",
      "content": "<p>For what I saw in the few games of rocket league I saw before this competition, if timed well at the kickoff it is possible to make a direct goal. <br>\nJust searching \"rocket league kickoff goal\" and filtering results to gifs I found this that shows a kickoff goal (ball in 0,0). <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6043604%2Fd06d06d5536c15e9fb1c8543e441556b%2FGloriousBowedHippopotamus-size_restricted.gif?generation=1665436806565565&amp;alt=media\" alt=\"\"><br>\n<a href=\"https://thumbs.gfycat.com/GloriousBowedHippopotamus-size_restricted.gif\" target=\"_blank\">source</a>.</p>\n<p>Even if there is no immediate goal, as stated by <a href=\"https://www.kaggle.com/maherelouahabi\" target=\"_blank\">@maherelouahabi</a> the first to touch the ball will likely have an advantage towards scoring a goal.</p>",
      "rawMarkdown": "For what I saw in the few games of rocket league I saw before this competition, if timed well at the kickoff it is possible to make a direct goal. \nJust searching \"rocket league kickoff goal\" and filtering results to gifs I found this that shows a kickoff goal (ball in 0,0). \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6043604%2Fd06d06d5536c15e9fb1c8543e441556b%2FGloriousBowedHippopotamus-size_restricted.gif?generation=1665436806565565&alt=media)\n[source](https://thumbs.gfycat.com/GloriousBowedHippopotamus-size_restricted.gif).\n\nEven if there is no immediate goal, as stated by @maherelouahabi the first to touch the ball will likely have an advantage towards scoring a goal.",
      "votes": null
    },
    {
      "id": "1981576",
      "postDate": "10/10/2022 21:51:40",
      "content": "<p><a href=\"https://www.kaggle.com/maherelouahabi\" target=\"_blank\">@maherelouahabi</a> As I understand it, you are essentially making the following calculations about posterior probability (using <code>train_0</code> dataset for simplicity)</p>\n<pre><code>display(train_0[(train_0.ball_pos_x.astype('int')==0)&amp;(train_0.ball_pos_y.astype('int')==0)]['team_A_scoring_within_10sec'].mean())\ndisplay(train_0[(train_0.ball_pos_x.astype('int')==0)&amp;(train_0.ball_pos_y.astype('int')==0)]['team_B_scoring_within_10sec'].mean())\n</code></pre>\n<pre><code>0.03609208558716562\n0.0436428083519486\n</code></pre>\n<p>Now if I arbitrarily take a different position, say \\((10,\\pm20)\\), I get better probabilities</p>\n<pre><code>display(train_0[(train_0.ball_pos_x.astype('int')==10)&amp;(train_0.ball_pos_y.astype('int')==20)]['team_A_scoring_within_10sec'].mean())\ndisplay(train_0[(train_0.ball_pos_x.astype('int')==10)&amp;(train_0.ball_pos_y.astype('int')==-20)]['team_B_scoring_within_10sec'].mean())\n</code></pre>\n<pre><code>0.07142857142857142\n0.09523809523809523\n</code></pre>",
      "rawMarkdown": "maherelouahabi As I understand it, you are essentially making the following calculations about posterior probability (using `train_0` dataset for simplicity)\n```\ndisplay(train_0[(train_0.ball_pos_x.astype('int')==0)&(train_0.ball_pos_y.astype('int')==0)]['team_A_scoring_within_10sec'].mean())\ndisplay(train_0[(train_0.ball_pos_x.astype('int')==0)&(train_0.ball_pos_y.astype('int')==0)]['team_B_scoring_within_10sec'].mean())\n```\n```\n0.03609208558716562\n0.0436428083519486\n```\nNow if I arbitrarily take a different position, say \\\\((10,\\pm20)\\\\), I get better probabilities\n```\ndisplay(train_0[(train_0.ball_pos_x.astype('int')==10)&(train_0.ball_pos_y.astype('int')==20)]['team_A_scoring_within_10sec'].mean())\ndisplay(train_0[(train_0.ball_pos_x.astype('int')==10)&(train_0.ball_pos_y.astype('int')==-20)]['team_B_scoring_within_10sec'].mean())\n```\n```\n0.07142857142857142\n0.09523809523809523\n```",
      "votes": null
    },
    {
      "id": "1981990",
      "postDate": "10/11/2022 07:12:12",
      "content": "<p>Dear Broccoli,</p>\n<p>You are calculating <strong>p(team_X_scoring_within_10sec=1/position)</strong> of team X, I am calculating <strong>p(position/team_X_scoring_within_10sec=1)</strong> of team X.</p>\n<p>Regards</p>",
      "rawMarkdown": "Dear Broccoli,\n\nYou are calculating **p(team_X_scoring_within_10sec=1/position)** of team X, I am calculating **p(position/team_X_scoring_within_10sec=1)** of team X.\n\nRegards",
      "votes": null
    },
    {
      "id": "1982009",
      "postDate": "10/11/2022 07:30:32",
      "content": "<p><a href=\"https://www.kaggle.com/maherelouahabi\" target=\"_blank\">@maherelouahabi</a> My calculations are for the posterior probability \\(p(y|x)\\) where \\(y\\) is a binary target variable and \\(x\\) is a categorical variable (e.g., the quantized \\(x\\)- and \\(y\\)-coordinates). In the case of a binary target variable, this is essentially the target encoding without smoothing. </p>\n<p>It looks like you were calculating \\(p(x|y=1)\\) and that would explain the spike at \\(x=(0,0)\\). This \"feature\" would not have any correlation with the target variable \\(y\\) for the obvious reason that it is calculated using only one class of \\(y\\).</p>",
      "rawMarkdown": "maherelouahabi My calculations are for the posterior probability \\\\(p(y|x)\\\\) where \\\\(y\\\\) is a binary target variable and \\\\(x\\\\) is a categorical variable (e.g., the quantized \\\\(x\\\\)- and \\\\(y\\\\)-coordinates). In the case of a binary target variable, this is essentially the target encoding without smoothing. \n\nIt looks like you were calculating \\\\(p(x|y=1)\\\\) and that would explain the spike at \\\\(x=(0,0)\\\\). This \"feature\" would not have any correlation with the target variable \\\\(y\\\\) for the obvious reason that it is calculated using only one class of \\\\(y\\\\).",
      "votes": null
    },
    {
      "id": "1982144",
      "postDate": "10/11/2022 09:30:24",
      "content": "<p>I also don't see the spike at 0,0 <a href=\"https://www.kaggle.com/code/paddykb/tps-2022-10-r-scoring-probability-surface/\" target=\"_blank\">fitting a smooth probabilty surface</a> (logistic regression using a tensor smooth on ball position)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F175917%2F6ee0aa75a172f7da206781c0aab1ac27%2Fsurface.png?generation=1665644869987636&amp;alt=media\" alt=\"\"><br>\nLeast chance of scoring when the ball is in our own goal :D</p>",
      "rawMarkdown": "I also don't see the spike at 0,0 [fitting a smooth probabilty surface](https://www.kaggle.com/code/paddykb/tps-2022-10-r-scoring-probability-surface/) (logistic regression using a tensor smooth on ball position)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F175917%2F6ee0aa75a172f7da206781c0aab1ac27%2Fsurface.png?generation=1665644869987636&alt=media)\nLeast chance of scoring when the ball is in our own goal :D",
      "votes": null
    },
    {
      "id": "1982265",
      "postDate": "10/11/2022 10:57:17",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/paddykb\" target=\"_blank\">@paddykb</a>,</p>\n<p>Thank you very much for point this out. The point (0, 0) is a high probability point only on kick-off or after scoring a goal for both teams. When calculating the probability, I have not disaggregated the temporal information of this point, causing that point to be a high probability point. During inference we do not have this temporal information so I think it is important to disaggregate this. I will correct it.</p>\n<p>Regards<br>\nMaher</p>",
      "rawMarkdown": "Dear @paddykb,\n\nThank you very much for point this out. The point (0, 0) is a high probability point only on kick-off or after scoring a goal for both teams. When calculating the probability, I have not disaggregated the temporal information of this point, causing that point to be a high probability point. During inference we do not have this temporal information so I think it is important to disaggregate this. I will correct it.\n\nRegards\nMaher",
      "votes": null
    },
    {
      "id": "1983190",
      "postDate": "10/11/2022 20:46:34",
      "content": "<p><a href=\"https://www.kaggle.com/paddykb\" target=\"_blank\">@paddykb</a> Instead of modeling the usual posterior probability \\(p(y=1|x)\\), OP calculates instead \\(p(x|y=1)\\) which basically asks, for all the ball trajectories that end up in the goal and take at most 10 sec, what is the distribution of the starting point. If the time limit was 1000 sec (or something essentially \\(\\infty\\)) I can see that the distribution would have a spike at \\((0,0)\\) simply because every trajectory starts from \\((0,0)\\). For 10 sec time limit, some trajectories still survive so it is likely there is still a spike at \\((0,0)\\). As I said, \\(p(x|y=1)\\) is not predictive of the variable \\(y\\).</p>",
      "rawMarkdown": "paddykb Instead of modeling the usual posterior probability \\\\(p(y=1|x)\\\\), OP calculates instead \\\\(p(x|y=1)\\\\) which basically asks, for all the ball trajectories that end up in the goal and take at most 10 sec, what is the distribution of the starting point. If the time limit was 1000 sec (or something essentially \\\\(\\infty\\\\)) I can see that the distribution would have a spike at \\\\((0,0)\\\\) simply because every trajectory starts from \\\\((0,0)\\\\). For 10 sec time limit, some trajectories still survive so it is likely there is still a spike at \\\\((0,0)\\\\). As I said, \\\\(p(x|y=1)\\\\) is not predictive of the variable \\\\(y\\\\).",
      "votes": null
    },
    {
      "id": "1983283",
      "postDate": "10/11/2022 23:43:57",
      "content": "<p>Great insights 🤯. Thanks for sharing! I will try to use these dictionaries as features</p>",
      "rawMarkdown": "Great insights 🤯. Thanks for sharing! I will try to use these dictionaries as features",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1979551,
      "author_name": "pietromaldini1",
      "author_url": "",
      "post_date": "10/09/2022 14:43:17",
      "content": "<p>Really an insightful post, I will definitely try to the dictionaries as features, I may even try and remove the real ball position from the features used in a model. I have the feeling that many models are not working to their full potential since the position of players and positions of balls are hard to use as they are.</p>\n<p>A really important aspect to notice is that the distance is not the only aspect that makes up the likelihood, also being or not at an edge of the field seems to contribute, another insight that may be used for FE. </p>\n<p>Thank you <a href=\"https://www.kaggle.com/maherelouahabi\" target=\"_blank\">@maherelouahabi</a> ! Keep kaggling and let's enjoy this month's TPS</p>",
      "votes": null,
      "replies": [
        {
          "id": 1979634,
          "author_name": "maherelouahabi",
          "author_url": "",
          "post_date": "10/09/2022 15:37:40",
          "content": "<p>Dear Pietro,</p>\n<p>I think one of the main problem with the positions is the high factorial cardinality of them. Approaches like this one, cook this high cardinality.<br>\nVery interesting what you say about the other factors that increase likelihood. I will think about it.</p>\n<p>Happy kaggling <a href=\"https://www.kaggle.com/pietromaldini1\" target=\"_blank\">@pietromaldini1</a> <br>\nRegards</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1979585,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "10/09/2022 15:03:05",
      "content": "<p>Great observations <a href=\"https://www.kaggle.com/maherelouahabi\" target=\"_blank\">@maherelouahabi</a>, thanks for sharing this information too!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1979628,
          "author_name": "maherelouahabi",
          "author_url": "",
          "post_date": "10/09/2022 15:31:26",
          "content": "<p>You are welcome. Happy kaggling <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1981261,
      "author_name": "pastorsoto",
      "author_url": "",
      "post_date": "10/10/2022 17:33:00",
      "content": "<p>Amazing insight! In soccer, there is a thing called Expected Goal that match the probability of scoring with an area of the field, some people even use the position of the ball to optimize this probabilities</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2224301%2Fda46c492db121d5bf0b14b0c17ea10b2%2Ffutbol_prob.jpg?generation=1665423082433728&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1981556,
          "author_name": "maherelouahabi",
          "author_url": "",
          "post_date": "10/10/2022 21:16:46",
          "content": "<p>Dear Pastor,</p>\n<p>Very interesting what you say. I didn't know that a similar concept exists in soccer. Thank you very much for the contribution.</p>\n<p>Happy kaggling <a href=\"https://www.kaggle.com/pastorsoto\" target=\"_blank\">@pastorsoto</a> <br>\nRegards</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1981529,
      "author_name": "siukeitin",
      "author_url": "",
      "post_date": "10/10/2022 20:38:37",
      "content": "<p>It seems that these features are related to target encoding the integer ball coordinates with respect to the two target variables. For example,</p>\n<pre><code>from category_encoders.target_encoder import TargetEncoder\n\ntrain['int_ball_pos']=train.ball_pos_x.astype('int').astype('str')+'_'+train.ball_pos_y.astype('int').astype('str')\nenc_A = TargetEncoder(return_df=False).fit(train[['int_ball_pos']],train.team_A_scoring_within_10sec)\nenc_B = TargetEncoder(return_df=False).fit(train[['int_ball_pos']],train.team_B_scoring_within_10sec)\ntrain['ball_pos_prob_A']=enc_A.transform(train[['int_ball_pos']])\ntrain['ball_pos_prob_B']=enc_B.transform(train[['int_ball_pos']])\ntrain.drop(['int_ball_pos'],axis=1,inplace=True)\n</code></pre>\n<p>Why does ball position \\((0,0)\\) have high probability of scoring within the next 10 sec for either team? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1981558,
          "author_name": "maherelouahabi",
          "author_url": "",
          "post_date": "10/10/2022 21:21:39",
          "content": "<p>Dear Brococoli,</p>\n<p>I was waiting for this quesiton. The main reason is because it is the point at which the match starts + the point at which they start after scoring a goal. If you look at any gameplay, as soon as the match starts, the players closest to the position of the ball, go as fast as possible to the ball. The team that arrives first, usually transmits a lot of inertia to the ball in the direction of the opponent's goal, often ending in a goal.</p>\n<p>Happy kaggling <a href=\"https://www.kaggle.com/siukeitin\" target=\"_blank\">@siukeitin</a> <br>\nRegards</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1981559,
          "author_name": "pietromaldini1",
          "author_url": "",
          "post_date": "10/10/2022 21:22:40",
          "content": "<p>For what I saw in the few games of rocket league I saw before this competition, if timed well at the kickoff it is possible to make a direct goal. <br>\nJust searching \"rocket league kickoff goal\" and filtering results to gifs I found this that shows a kickoff goal (ball in 0,0). <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6043604%2Fd06d06d5536c15e9fb1c8543e441556b%2FGloriousBowedHippopotamus-size_restricted.gif?generation=1665436806565565&amp;alt=media\" alt=\"\"><br>\n<a href=\"https://thumbs.gfycat.com/GloriousBowedHippopotamus-size_restricted.gif\" target=\"_blank\">source</a>.</p>\n<p>Even if there is no immediate goal, as stated by <a href=\"https://www.kaggle.com/maherelouahabi\" target=\"_blank\">@maherelouahabi</a> the first to touch the ball will likely have an advantage towards scoring a goal.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1981576,
          "author_name": "siukeitin",
          "author_url": "",
          "post_date": "10/10/2022 21:51:40",
          "content": "<p><a href=\"https://www.kaggle.com/maherelouahabi\" target=\"_blank\">@maherelouahabi</a> As I understand it, you are essentially making the following calculations about posterior probability (using <code>train_0</code> dataset for simplicity)</p>\n<pre><code>display(train_0[(train_0.ball_pos_x.astype('int')==0)&amp;(train_0.ball_pos_y.astype('int')==0)]['team_A_scoring_within_10sec'].mean())\ndisplay(train_0[(train_0.ball_pos_x.astype('int')==0)&amp;(train_0.ball_pos_y.astype('int')==0)]['team_B_scoring_within_10sec'].mean())\n</code></pre>\n<pre><code>0.03609208558716562\n0.0436428083519486\n</code></pre>\n<p>Now if I arbitrarily take a different position, say \\((10,\\pm20)\\), I get better probabilities</p>\n<pre><code>display(train_0[(train_0.ball_pos_x.astype('int')==10)&amp;(train_0.ball_pos_y.astype('int')==20)]['team_A_scoring_within_10sec'].mean())\ndisplay(train_0[(train_0.ball_pos_x.astype('int')==10)&amp;(train_0.ball_pos_y.astype('int')==-20)]['team_B_scoring_within_10sec'].mean())\n</code></pre>\n<pre><code>0.07142857142857142\n0.09523809523809523\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1981990,
          "author_name": "maherelouahabi",
          "author_url": "",
          "post_date": "10/11/2022 07:12:12",
          "content": "<p>Dear Broccoli,</p>\n<p>You are calculating <strong>p(team_X_scoring_within_10sec=1/position)</strong> of team X, I am calculating <strong>p(position/team_X_scoring_within_10sec=1)</strong> of team X.</p>\n<p>Regards</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1982009,
          "author_name": "siukeitin",
          "author_url": "",
          "post_date": "10/11/2022 07:30:32",
          "content": "<p><a href=\"https://www.kaggle.com/maherelouahabi\" target=\"_blank\">@maherelouahabi</a> My calculations are for the posterior probability \\(p(y|x)\\) where \\(y\\) is a binary target variable and \\(x\\) is a categorical variable (e.g., the quantized \\(x\\)- and \\(y\\)-coordinates). In the case of a binary target variable, this is essentially the target encoding without smoothing. </p>\n<p>It looks like you were calculating \\(p(x|y=1)\\) and that would explain the spike at \\(x=(0,0)\\). This \"feature\" would not have any correlation with the target variable \\(y\\) for the obvious reason that it is calculated using only one class of \\(y\\).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1982144,
          "author_name": "paddykb",
          "author_url": "",
          "post_date": "10/11/2022 09:30:24",
          "content": "<p>I also don't see the spike at 0,0 <a href=\"https://www.kaggle.com/code/paddykb/tps-2022-10-r-scoring-probability-surface/\" target=\"_blank\">fitting a smooth probabilty surface</a> (logistic regression using a tensor smooth on ball position)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F175917%2F6ee0aa75a172f7da206781c0aab1ac27%2Fsurface.png?generation=1665644869987636&amp;alt=media\" alt=\"\"><br>\nLeast chance of scoring when the ball is in our own goal :D</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1982265,
          "author_name": "maherelouahabi",
          "author_url": "",
          "post_date": "10/11/2022 10:57:17",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/paddykb\" target=\"_blank\">@paddykb</a>,</p>\n<p>Thank you very much for point this out. The point (0, 0) is a high probability point only on kick-off or after scoring a goal for both teams. When calculating the probability, I have not disaggregated the temporal information of this point, causing that point to be a high probability point. During inference we do not have this temporal information so I think it is important to disaggregate this. I will correct it.</p>\n<p>Regards<br>\nMaher</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1983190,
          "author_name": "siukeitin",
          "author_url": "",
          "post_date": "10/11/2022 20:46:34",
          "content": "<p><a href=\"https://www.kaggle.com/paddykb\" target=\"_blank\">@paddykb</a> Instead of modeling the usual posterior probability \\(p(y=1|x)\\), OP calculates instead \\(p(x|y=1)\\) which basically asks, for all the ball trajectories that end up in the goal and take at most 10 sec, what is the distribution of the starting point. If the time limit was 1000 sec (or something essentially \\(\\infty\\)) I can see that the distribution would have a spike at \\((0,0)\\) simply because every trajectory starts from \\((0,0)\\). For 10 sec time limit, some trajectories still survive so it is likely there is still a spike at \\((0,0)\\). As I said, \\(p(x|y=1)\\) is not predictive of the variable \\(y\\).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1983283,
      "author_name": "mrgabrielblins",
      "author_url": "",
      "post_date": "10/11/2022 23:43:57",
      "content": "<p>Great insights 🤯. Thanks for sharing! I will try to use these dictionaries as features</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1979475": "As many of you will have deduced, the closer the ball is to the opponent's goal, the greater the probability of scoring a goal. \n\nA logical question that may arise is: \n**What does this probability distribution look like if we only look at the position of the ball and completely ignore the location of the player?** \n\nWell, if we analyse all the training data we get a distribution for each team that looks something like this:\n\nTeam A:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4696479%2Fff0b3602d1033659c2ce68ae341b3865%2F__results___3_2.png?generation=1665322117461398&alt=media)\n\nTeam B:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4696479%2F0bcc2e04ac384e124bb5f533585b3eec%2F__results___3_4.png?generation=1665322163197539&alt=media)\n\nAs we can see, apart from the goal area there are two privileged positions of high goal probability: the kick-off zone and the touchline at the back of the pitch.\n\nI have published these two dictionaries so that you can use this information as an additional feature:\nhttps://www.kaggle.com/datasets/maherelouahabi/court-probability-distrubution/\n\nHow to use it:\n```\nimport pickle\nimport pandas as pd\n\ndf = pd.read_csv('...')\n\ndef open_pkl(name):\n    with open(f'{name}.pkl', 'rb') as f:\n        return pickle.load(f)\n\ndict_A = open_pkl('../input/court-probability/dict_A')\n\ndict_B = open_pkl('../input/court-probability/dict_B')\n\ndf['ball_pos_prob_A'] = df.apply(lambda x: dict_A.get(f'{(round(x.ball_pos_y),round(x.ball_pos_x))}',0), axis=1)\ndf['ball_pos_prob_B'] = df.apply(lambda x: dict_B.get(f'{(round(x.ball_pos_y),round(x.ball_pos_x))}',0), axis=1)\n```\n\n**(UPDATE)**\nI have corrected the distribution calculation and updated both the dictionaries and the plots (Thank you @siukeitin). I leave the previous plots so that anyone can learn from my mistakes. \n\nTeam A:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4696479%2F89fefdbc1a68ef0c2dd824ad1c59c3f0%2FA.png?generation=1665569798239607&alt=media)\n\nTeam B:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4696479%2F32625d5a599b2208e2c8dbada98c0ba6%2FB.png?generation=1665569813790580&alt=media)",
    "1979551": "Really an insightful post, I will definitely try to the dictionaries as features, I may even try and remove the real ball position from the features used in a model. I have the feeling that many models are not working to their full potential since the position of players and positions of balls are hard to use as they are.\n\nA really important aspect to notice is that the distance is not the only aspect that makes up the likelihood, also being or not at an edge of the field seems to contribute, another insight that may be used for FE. \n\nThank you @maherelouahabi ! Keep kaggling and let's enjoy this month's TPS",
    "1979585": "Great observations @maherelouahabi, thanks for sharing this information too!",
    "1979628": "You are welcome. Happy kaggling @ravi20076",
    "1979634": "Dear Pietro,\n\nI think one of the main problem with the positions is the high factorial cardinality of them. Approaches like this one, cook this high cardinality.\nVery interesting what you say about the other factors that increase likelihood. I will think about it.\n\nHappy kaggling @pietromaldini1 \nRegards",
    "1981261": "Amazing insight! In soccer, there is a thing called Expected Goal that match the probability of scoring with an area of the field, some people even use the position of the ball to optimize this probabilities\n\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2224301%2Fda46c492db121d5bf0b14b0c17ea10b2%2Ffutbol_prob.jpg?generation=1665423082433728&alt=media)",
    "1981529": "It seems that these features are related to target encoding the integer ball coordinates with respect to the two target variables. For example,\n```\nfrom category_encoders.target_encoder import TargetEncoder\n\ntrain['int_ball_pos']=train.ball_pos_x.astype('int').astype('str')+'_'+train.ball_pos_y.astype('int').astype('str')\nenc_A = TargetEncoder(return_df=False).fit(train[['int_ball_pos']],train.team_A_scoring_within_10sec)\nenc_B = TargetEncoder(return_df=False).fit(train[['int_ball_pos']],train.team_B_scoring_within_10sec)\ntrain['ball_pos_prob_A']=enc_A.transform(train[['int_ball_pos']])\ntrain['ball_pos_prob_B']=enc_B.transform(train[['int_ball_pos']])\ntrain.drop(['int_ball_pos'],axis=1,inplace=True)\n```\nWhy does ball position \\\\((0,0)\\\\) have high probability of scoring within the next 10 sec for either team?",
    "1981556": "Dear Pastor,\n\nVery interesting what you say. I didn't know that a similar concept exists in soccer. Thank you very much for the contribution.\n\nHappy kaggling @pastorsoto \nRegards",
    "1981558": "Dear Brococoli,\n\nI was waiting for this quesiton. The main reason is because it is the point at which the match starts + the point at which they start after scoring a goal. If you look at any gameplay, as soon as the match starts, the players closest to the position of the ball, go as fast as possible to the ball. The team that arrives first, usually transmits a lot of inertia to the ball in the direction of the opponent's goal, often ending in a goal.\n\nHappy kaggling @siukeitin \nRegards",
    "1981559": "For what I saw in the few games of rocket league I saw before this competition, if timed well at the kickoff it is possible to make a direct goal. \nJust searching \"rocket league kickoff goal\" and filtering results to gifs I found this that shows a kickoff goal (ball in 0,0). \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6043604%2Fd06d06d5536c15e9fb1c8543e441556b%2FGloriousBowedHippopotamus-size_restricted.gif?generation=1665436806565565&alt=media)\n[source](https://thumbs.gfycat.com/GloriousBowedHippopotamus-size_restricted.gif).\n\nEven if there is no immediate goal, as stated by @maherelouahabi the first to touch the ball will likely have an advantage towards scoring a goal.",
    "1981576": "maherelouahabi As I understand it, you are essentially making the following calculations about posterior probability (using `train_0` dataset for simplicity)\n```\ndisplay(train_0[(train_0.ball_pos_x.astype('int')==0)&(train_0.ball_pos_y.astype('int')==0)]['team_A_scoring_within_10sec'].mean())\ndisplay(train_0[(train_0.ball_pos_x.astype('int')==0)&(train_0.ball_pos_y.astype('int')==0)]['team_B_scoring_within_10sec'].mean())\n```\n```\n0.03609208558716562\n0.0436428083519486\n```\nNow if I arbitrarily take a different position, say \\\\((10,\\pm20)\\\\), I get better probabilities\n```\ndisplay(train_0[(train_0.ball_pos_x.astype('int')==10)&(train_0.ball_pos_y.astype('int')==20)]['team_A_scoring_within_10sec'].mean())\ndisplay(train_0[(train_0.ball_pos_x.astype('int')==10)&(train_0.ball_pos_y.astype('int')==-20)]['team_B_scoring_within_10sec'].mean())\n```\n```\n0.07142857142857142\n0.09523809523809523\n```",
    "1981990": "Dear Broccoli,\n\nYou are calculating **p(team_X_scoring_within_10sec=1/position)** of team X, I am calculating **p(position/team_X_scoring_within_10sec=1)** of team X.\n\nRegards",
    "1982009": "maherelouahabi My calculations are for the posterior probability \\\\(p(y|x)\\\\) where \\\\(y\\\\) is a binary target variable and \\\\(x\\\\) is a categorical variable (e.g., the quantized \\\\(x\\\\)- and \\\\(y\\\\)-coordinates). In the case of a binary target variable, this is essentially the target encoding without smoothing. \n\nIt looks like you were calculating \\\\(p(x|y=1)\\\\) and that would explain the spike at \\\\(x=(0,0)\\\\). This \"feature\" would not have any correlation with the target variable \\\\(y\\\\) for the obvious reason that it is calculated using only one class of \\\\(y\\\\).",
    "1982144": "I also don't see the spike at 0,0 [fitting a smooth probabilty surface](https://www.kaggle.com/code/paddykb/tps-2022-10-r-scoring-probability-surface/) (logistic regression using a tensor smooth on ball position)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F175917%2F6ee0aa75a172f7da206781c0aab1ac27%2Fsurface.png?generation=1665644869987636&alt=media)\nLeast chance of scoring when the ball is in our own goal :D",
    "1982265": "Dear @paddykb,\n\nThank you very much for point this out. The point (0, 0) is a high probability point only on kick-off or after scoring a goal for both teams. When calculating the probability, I have not disaggregated the temporal information of this point, causing that point to be a high probability point. During inference we do not have this temporal information so I think it is important to disaggregate this. I will correct it.\n\nRegards\nMaher",
    "1983190": "paddykb Instead of modeling the usual posterior probability \\\\(p(y=1|x)\\\\), OP calculates instead \\\\(p(x|y=1)\\\\) which basically asks, for all the ball trajectories that end up in the goal and take at most 10 sec, what is the distribution of the starting point. If the time limit was 1000 sec (or something essentially \\\\(\\infty\\\\)) I can see that the distribution would have a spike at \\\\((0,0)\\\\) simply because every trajectory starts from \\\\((0,0)\\\\). For 10 sec time limit, some trajectories still survive so it is likely there is still a spike at \\\\((0,0)\\\\). As I said, \\\\(p(x|y=1)\\\\) is not predictive of the variable \\\\(y\\\\).",
    "1983283": "Great insights 🤯. Thanks for sharing! I will try to use these dictionaries as features"
  },
  "source": "meta"
}