{
  "id": 357577,
  "title": "Data augmentation - A possible approach",
  "url": "/competitions/tabular-playground-series-oct-2022/discussion/357577",
  "author_name": "",
  "post_date": "2022-10-04T19:32:22.532007100Z",
  "votes": 11,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I was thinking the last days about a possible way to perform data augmentation.<br>\nI thought about several things but the only one that I felt could help is changing player order.</p>\n<p>You could exchange player 1 and 2 (position/velocity and boost) and the model should predict the same outcome. </p>\n<p>The same goes with exchanging teams, but this may be trickier since in this case it would be better to also rotate the field, since players in team A starts always in the same half of the field and players in field B starts in the other half.</p>\n<p>So we can think that a model that correctly captures the problem should be invariant to certain permutations of players. Those permutation that keep the players in the same team can be used to augment the data.</p>\n<p>This idea should hold for this problem.</p>\n<p>When I come back here on kaggle I read the discussions left in the last days.<br>\nI noticed this <a href=\"https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/357339\" target=\"_blank\">topic</a> that suggests that the ordering of players is not random for this dataset.</p>\n<p>The implication of this topic:</p>\n<ol>\n<li>Permutation based augmentation may reduce the effectiveness of the model</li>\n<li>While probably reducing the effectiveness of the model on this dataset te use of permutation based augmentation, the generated model should be more resistant to an analogue dataset where the ordering of players is chosen at random, so low performing players should not be more likely to be in team B.</li>\n<li>We have still a lot to find out on the dataset. </li>\n<li>Lot's of fun waiting for us all participating</li>\n</ol>\n<p>Keep kaggling and enjoy this month's TPS.</p>",
  "messages": [
    {
      "id": "1971857",
      "postDate": "10/04/2022 19:32:22",
      "content": "<p>I was thinking the last days about a possible way to perform data augmentation.<br>\nI thought about several things but the only one that I felt could help is changing player order.</p>\n<p>You could exchange player 1 and 2 (position/velocity and boost) and the model should predict the same outcome. </p>\n<p>The same goes with exchanging teams, but this may be trickier since in this case it would be better to also rotate the field, since players in team A starts always in the same half of the field and players in field B starts in the other half.</p>\n<p>So we can think that a model that correctly captures the problem should be invariant to certain permutations of players. Those permutation that keep the players in the same team can be used to augment the data.</p>\n<p>This idea should hold for this problem.</p>\n<p>When I come back here on kaggle I read the discussions left in the last days.<br>\nI noticed this <a href=\"https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/357339\" target=\"_blank\">topic</a> that suggests that the ordering of players is not random for this dataset.</p>\n<p>The implication of this topic:</p>\n<ol>\n<li>Permutation based augmentation may reduce the effectiveness of the model</li>\n<li>While probably reducing the effectiveness of the model on this dataset te use of permutation based augmentation, the generated model should be more resistant to an analogue dataset where the ordering of players is chosen at random, so low performing players should not be more likely to be in team B.</li>\n<li>We have still a lot to find out on the dataset. </li>\n<li>Lot's of fun waiting for us all participating</li>\n</ol>\n<p>Keep kaggling and enjoy this month's TPS.</p>",
      "rawMarkdown": "I was thinking the last days about a possible way to perform data augmentation.\nI thought about several things but the only one that I felt could help is changing player order.\n\nYou could exchange player 1 and 2 (position/velocity and boost) and the model should predict the same outcome. \n\nThe same goes with exchanging teams, but this may be trickier since in this case it would be better to also rotate the field, since players in team A starts always in the same half of the field and players in field B starts in the other half.\n\nSo we can think that a model that correctly captures the problem should be invariant to certain permutations of players. Those permutation that keep the players in the same team can be used to augment the data.\n\nThis idea should hold for this problem.\n\nWhen I come back here on kaggle I read the discussions left in the last days.\nI noticed this [topic](https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/357339) that suggests that the ordering of players is not random for this dataset.\n\nThe implication of this topic:\n1. Permutation based augmentation may reduce the effectiveness of the model\n2. While probably reducing the effectiveness of the model on this dataset te use of permutation based augmentation, the generated model should be more resistant to an analogue dataset where the ordering of players is chosen at random, so low performing players should not be more likely to be in team B.\n3. We have still a lot to find out on the dataset. \n4. Lot's of fun waiting for us all participating\n\nKeep kaggling and enjoy this month's TPS.",
      "votes": null
    },
    {
      "id": "1972513",
      "postDate": "10/05/2022 07:33:25",
      "content": "<p>This looks like an excellent idea.</p>\n<p>I was looking at building a NN and augmenting the batches with random permutations from the 144 possible ones:</p>\n<ul>\n<li>Six ways to order team A players.</li>\n<li>Six ways to order team B players</li>\n<li>flip teams A and B</li>\n<li>horizontal reflection</li>\n</ul>\n<p>Since we have event order in the train set, it would be interesting to model successive positions. I.e. use the current positions as input and the next position of players and ball as the output. (strip off the last layer and use it as an embedding).</p>",
      "rawMarkdown": "This looks like an excellent idea.\n\nI was looking at building a NN and augmenting the batches with random permutations from the 144 possible ones:\n\n- Six ways to order team A players.\n- Six ways to order team B players\n- flip teams A and B\n- horizontal reflection\n\nSince we have event order in the train set, it would be interesting to model successive positions. I.e. use the current positions as input and the next position of players and ball as the output. (strip off the last layer and use it as an embedding).",
      "votes": null
    },
    {
      "id": "1972542",
      "postDate": "10/05/2022 07:46:24",
      "content": "<p>I feel that your idea of predicting future positions of players and ball is great. If you can model accurately the game and get low error you could even simulate the game up to 10 seconds in the future or until the ball is within goal position, at that point you should find out quite easily which team scored.</p>",
      "rawMarkdown": "I feel that your idea of predicting future positions of players and ball is great. If you can model accurately the game and get low error you could even simulate the game up to 10 seconds in the future or until the ball is within goal position, at that point you should find out quite easily which team scored.",
      "votes": null
    },
    {
      "id": "1972912",
      "postDate": "10/05/2022 11:40:25",
      "content": "<p>You may be able to break the player symmetry by sorting the players dynamically in some way, e.g., relative to distance to the current location of the ball. In other words, instead of having features pertaining to \\(p_0,p_1,p_2\\) that are designated to fixed players, you would have features pertaining to say \\(q_0,q_1,q_2\\) where \\(q_0\\in\\{p_0,p_1,p_2\\}\\) is the player closest to the ball at that time instant, \\(q_1\\) is next etc.</p>",
      "rawMarkdown": "You may be able to break the player symmetry by sorting the players dynamically in some way, e.g., relative to distance to the current location of the ball. In other words, instead of having features pertaining to \\\\(p_0,p_1,p_2\\\\) that are designated to fixed players, you would have features pertaining to say \\\\(q_0,q_1,q_2\\\\) where \\\\(q_0\\in\\\\{p_0,p_1,p_2\\\\}\\\\) is the player closest to the ball at that time instant, \\\\(q_1\\\\) is next etc.",
      "votes": null
    },
    {
      "id": "1972939",
      "postDate": "10/05/2022 11:50:57",
      "content": "<p>Neat. That's a nice representation. (Better than the huge sparse array that I've considered).</p>",
      "rawMarkdown": "Neat. That's a nice representation. (Better than the huge sparse array that I've considered).",
      "votes": null
    },
    {
      "id": "1973343",
      "postDate": "10/05/2022 15:23:09",
      "content": "<p>I had similar ideas :)<br>\nI started to experiment with exchanging the teams in my <a href=\"https://www.kaggle.com/code/spyrow/playground-oct-2022-lgmbclassifier/notebook\" target=\"_blank\">notebook</a>. (It is not very well documented right now). <br>\nI switched the teams in all datasets. Therefore i have doubled the training data. Furthermore i only need one model to predict if team a or team B will score a goal. I only need to switch the teams on the test data before predicting Team B.<br>\nI still need to check if i made any coding or logical errors. In addition to that i am also curious about this <a href=\"https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/357339\" target=\"_blank\">topic</a>, because my combined model is also a little bit worse then my approach with two different models on the test data.</p>",
      "rawMarkdown": "I had similar ideas :)\nI started to experiment with exchanging the teams in my [notebook](https://www.kaggle.com/code/spyrow/playground-oct-2022-lgmbclassifier/notebook). (It is not very well documented right now). \nI switched the teams in all datasets. Therefore i have doubled the training data. Furthermore i only need one model to predict if team a or team B will score a goal. I only need to switch the teams on the test data before predicting Team B.\nI still need to check if i made any coding or logical errors. In addition to that i am also curious about this [topic](https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/357339), because my combined model is also a little bit worse then my approach with two different models on the test data.",
      "votes": null
    },
    {
      "id": "1973364",
      "postDate": "10/05/2022 15:31:51",
      "content": "<p>Nice. I made a start testing <a href=\"https://www.kaggle.com/code/paddykb/tps-2022-10-fastai\" target=\"_blank\">permuting players</a>. I'm going to steal your mirroring code, and add an horizontal flip. Augmenting is a little easier using a NN because we can apply transformations on-the-fly to batches.</p>",
      "rawMarkdown": "Nice. I made a start testing [permuting players](https://www.kaggle.com/code/paddykb/tps-2022-10-fastai). I'm going to steal your mirroring code, and add an horizontal flip. Augmenting is a little easier using a NN because we can apply transformations on-the-fly to batches.",
      "votes": null
    },
    {
      "id": "1973736",
      "postDate": "10/05/2022 19:57:47",
      "content": "<p><a href=\"https://www.kaggle.com/spyrow\" target=\"_blank\">@spyrow</a> Interesting idea to use one binary classifier to predict both probabilities on \"flipped\" data. I don't see any mechanism in this approach to prevent the two probabilities from adding up greater than one (which should not happen with mutually exclusive outcomes), although the competition scoring metric does not forbid that. </p>",
      "rawMarkdown": "spyrow Interesting idea to use one binary classifier to predict both probabilities on \"flipped\" data. I don't see any mechanism in this approach to prevent the two probabilities from adding up greater than one (which should not happen with mutually exclusive outcomes), although the competition scoring metric does not forbid that.",
      "votes": null
    },
    {
      "id": "1973777",
      "postDate": "10/05/2022 20:34:13",
      "content": "<p>If i have some time i could try a classifier with three outputs on the combined data and compare the result. Right know i need to figure out some memory issues with my LGBM model. Maybe i will also try a NN on this.</p>",
      "rawMarkdown": "If i have some time i could try a classifier with three outputs on the combined data and compare the result. Right know i need to figure out some memory issues with my LGBM model. Maybe i will also try a NN on this.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1972513,
      "author_name": "paddykb",
      "author_url": "",
      "post_date": "10/05/2022 07:33:25",
      "content": "<p>This looks like an excellent idea.</p>\n<p>I was looking at building a NN and augmenting the batches with random permutations from the 144 possible ones:</p>\n<ul>\n<li>Six ways to order team A players.</li>\n<li>Six ways to order team B players</li>\n<li>flip teams A and B</li>\n<li>horizontal reflection</li>\n</ul>\n<p>Since we have event order in the train set, it would be interesting to model successive positions. I.e. use the current positions as input and the next position of players and ball as the output. (strip off the last layer and use it as an embedding).</p>",
      "votes": null,
      "replies": [
        {
          "id": 1972542,
          "author_name": "pietromaldini1",
          "author_url": "",
          "post_date": "10/05/2022 07:46:24",
          "content": "<p>I feel that your idea of predicting future positions of players and ball is great. If you can model accurately the game and get low error you could even simulate the game up to 10 seconds in the future or until the ball is within goal position, at that point you should find out quite easily which team scored.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1972912,
          "author_name": "siukeitin",
          "author_url": "",
          "post_date": "10/05/2022 11:40:25",
          "content": "<p>You may be able to break the player symmetry by sorting the players dynamically in some way, e.g., relative to distance to the current location of the ball. In other words, instead of having features pertaining to \\(p_0,p_1,p_2\\) that are designated to fixed players, you would have features pertaining to say \\(q_0,q_1,q_2\\) where \\(q_0\\in\\{p_0,p_1,p_2\\}\\) is the player closest to the ball at that time instant, \\(q_1\\) is next etc.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1972939,
          "author_name": "paddykb",
          "author_url": "",
          "post_date": "10/05/2022 11:50:57",
          "content": "<p>Neat. That's a nice representation. (Better than the huge sparse array that I've considered).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1973343,
      "author_name": "spyrow",
      "author_url": "",
      "post_date": "10/05/2022 15:23:09",
      "content": "<p>I had similar ideas :)<br>\nI started to experiment with exchanging the teams in my <a href=\"https://www.kaggle.com/code/spyrow/playground-oct-2022-lgmbclassifier/notebook\" target=\"_blank\">notebook</a>. (It is not very well documented right now). <br>\nI switched the teams in all datasets. Therefore i have doubled the training data. Furthermore i only need one model to predict if team a or team B will score a goal. I only need to switch the teams on the test data before predicting Team B.<br>\nI still need to check if i made any coding or logical errors. In addition to that i am also curious about this <a href=\"https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/357339\" target=\"_blank\">topic</a>, because my combined model is also a little bit worse then my approach with two different models on the test data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1973364,
          "author_name": "paddykb",
          "author_url": "",
          "post_date": "10/05/2022 15:31:51",
          "content": "<p>Nice. I made a start testing <a href=\"https://www.kaggle.com/code/paddykb/tps-2022-10-fastai\" target=\"_blank\">permuting players</a>. I'm going to steal your mirroring code, and add an horizontal flip. Augmenting is a little easier using a NN because we can apply transformations on-the-fly to batches.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1973736,
          "author_name": "siukeitin",
          "author_url": "",
          "post_date": "10/05/2022 19:57:47",
          "content": "<p><a href=\"https://www.kaggle.com/spyrow\" target=\"_blank\">@spyrow</a> Interesting idea to use one binary classifier to predict both probabilities on \"flipped\" data. I don't see any mechanism in this approach to prevent the two probabilities from adding up greater than one (which should not happen with mutually exclusive outcomes), although the competition scoring metric does not forbid that. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1973777,
          "author_name": "spyrow",
          "author_url": "",
          "post_date": "10/05/2022 20:34:13",
          "content": "<p>If i have some time i could try a classifier with three outputs on the combined data and compare the result. Right know i need to figure out some memory issues with my LGBM model. Maybe i will also try a NN on this.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1971857": "I was thinking the last days about a possible way to perform data augmentation.\nI thought about several things but the only one that I felt could help is changing player order.\n\nYou could exchange player 1 and 2 (position/velocity and boost) and the model should predict the same outcome. \n\nThe same goes with exchanging teams, but this may be trickier since in this case it would be better to also rotate the field, since players in team A starts always in the same half of the field and players in field B starts in the other half.\n\nSo we can think that a model that correctly captures the problem should be invariant to certain permutations of players. Those permutation that keep the players in the same team can be used to augment the data.\n\nThis idea should hold for this problem.\n\nWhen I come back here on kaggle I read the discussions left in the last days.\nI noticed this [topic](https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/357339) that suggests that the ordering of players is not random for this dataset.\n\nThe implication of this topic:\n1. Permutation based augmentation may reduce the effectiveness of the model\n2. While probably reducing the effectiveness of the model on this dataset te use of permutation based augmentation, the generated model should be more resistant to an analogue dataset where the ordering of players is chosen at random, so low performing players should not be more likely to be in team B.\n3. We have still a lot to find out on the dataset. \n4. Lot's of fun waiting for us all participating\n\nKeep kaggling and enjoy this month's TPS.",
    "1972513": "This looks like an excellent idea.\n\nI was looking at building a NN and augmenting the batches with random permutations from the 144 possible ones:\n\n- Six ways to order team A players.\n- Six ways to order team B players\n- flip teams A and B\n- horizontal reflection\n\nSince we have event order in the train set, it would be interesting to model successive positions. I.e. use the current positions as input and the next position of players and ball as the output. (strip off the last layer and use it as an embedding).",
    "1972542": "I feel that your idea of predicting future positions of players and ball is great. If you can model accurately the game and get low error you could even simulate the game up to 10 seconds in the future or until the ball is within goal position, at that point you should find out quite easily which team scored.",
    "1972912": "You may be able to break the player symmetry by sorting the players dynamically in some way, e.g., relative to distance to the current location of the ball. In other words, instead of having features pertaining to \\\\(p_0,p_1,p_2\\\\) that are designated to fixed players, you would have features pertaining to say \\\\(q_0,q_1,q_2\\\\) where \\\\(q_0\\in\\\\{p_0,p_1,p_2\\\\}\\\\) is the player closest to the ball at that time instant, \\\\(q_1\\\\) is next etc.",
    "1972939": "Neat. That's a nice representation. (Better than the huge sparse array that I've considered).",
    "1973343": "I had similar ideas :)\nI started to experiment with exchanging the teams in my [notebook](https://www.kaggle.com/code/spyrow/playground-oct-2022-lgmbclassifier/notebook). (It is not very well documented right now). \nI switched the teams in all datasets. Therefore i have doubled the training data. Furthermore i only need one model to predict if team a or team B will score a goal. I only need to switch the teams on the test data before predicting Team B.\nI still need to check if i made any coding or logical errors. In addition to that i am also curious about this [topic](https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/357339), because my combined model is also a little bit worse then my approach with two different models on the test data.",
    "1973364": "Nice. I made a start testing [permuting players](https://www.kaggle.com/code/paddykb/tps-2022-10-fastai). I'm going to steal your mirroring code, and add an horizontal flip. Augmenting is a little easier using a NN because we can apply transformations on-the-fly to batches.",
    "1973736": "spyrow Interesting idea to use one binary classifier to predict both probabilities on \"flipped\" data. I don't see any mechanism in this approach to prevent the two probabilities from adding up greater than one (which should not happen with mutually exclusive outcomes), although the competition scoring metric does not forbid that.",
    "1973777": "If i have some time i could try a classifier with three outputs on the combined data and compare the result. Right know i need to figure out some memory issues with my LGBM model. Maybe i will also try a NN on this."
  },
  "source": "meta"
}