{
  "id": 359502,
  "title": "Test time data augmentation",
  "url": "/competitions/tabular-playground-series-oct-2022/discussion/359502",
  "author_name": "Alexander Ryzhkov",
  "post_date": "2022-10-12T10:33:38.155000",
  "votes": 14,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi colleagues,</p>\n<p>First of all, I would like to mention the great appreciation for <a href=\"https://www.kaggle.com/paddykb\" target=\"_blank\">@paddykb</a> and <a href=\"https://www.kaggle.com/code/paddykb/tps-2022-10-fastai\" target=\"_blank\">his great kernel</a> - it is really awesome and you should take a look at it if you start working on this competition. </p>\n<p>When you are working on this kind of data, you will find some useful training data augmentations - shuffling players inside team, changing the teams in the mirror way (with changing targets as well) etc. But you can also do this for the test data - this is usually called Test Time Augmentation (TTA for short). You can do that because you know that the real target of the test data doesn't change if you do this kind of augmentation (all of them without mirror) so you can predict for the augmented test data several times and mean the received predictions.</p>\n<p>For the neural networks, it's also will be better if you use multi-start technique - run the network training several times and also calculate the mean predictions for the runs.</p>\n<p>Both of these techniques helped me to improve the <a href=\"https://www.kaggle.com/paddykb\" target=\"_blank\">@paddykb</a>'s kernel and now you can take a look of their implementation in my kernel (which is the cloned modified version of <a href=\"https://www.kaggle.com/code/paddykb/tps-2022-10-fastai\" target=\"_blank\">this kernel</a>: <a href=\"https://www.kaggle.com/alexryzhkov/tps-2022-10-fastai-with-multistart-and-tta\" target=\"_blank\">link to kernel with TTA and multistart</a></p>\n<p>Hope this can help you to improve your scores.</p>\n<p>Alex</p>",
  "messages": [
    {
      "id": 1983906,
      "postDate": "2022-10-12T10:33:38.157Z",
      "content": "<p>Hi colleagues,</p>\n<p>First of all, I would like to mention the great appreciation for <a href=\"https://www.kaggle.com/paddykb\" target=\"_blank\">@paddykb</a> and <a href=\"https://www.kaggle.com/code/paddykb/tps-2022-10-fastai\" target=\"_blank\">his great kernel</a> - it is really awesome and you should take a look at it if you start working on this competition. </p>\n<p>When you are working on this kind of data, you will find some useful training data augmentations - shuffling players inside team, changing the teams in the mirror way (with changing targets as well) etc. But you can also do this for the test data - this is usually called Test Time Augmentation (TTA for short). You can do that because you know that the real target of the test data doesn't change if you do this kind of augmentation (all of them without mirror) so you can predict for the augmented test data several times and mean the received predictions.</p>\n<p>For the neural networks, it's also will be better if you use multi-start technique - run the network training several times and also calculate the mean predictions for the runs.</p>\n<p>Both of these techniques helped me to improve the <a href=\"https://www.kaggle.com/paddykb\" target=\"_blank\">@paddykb</a>'s kernel and now you can take a look of their implementation in my kernel (which is the cloned modified version of <a href=\"https://www.kaggle.com/code/paddykb/tps-2022-10-fastai\" target=\"_blank\">this kernel</a>: <a href=\"https://www.kaggle.com/alexryzhkov/tps-2022-10-fastai-with-multistart-and-tta\" target=\"_blank\">link to kernel with TTA and multistart</a></p>\n<p>Hope this can help you to improve your scores.</p>\n<p>Alex</p>",
      "rawMarkdown": "Hi colleagues,\n\nFirst of all, I would like to mention the great appreciation for @paddykb and [his great kernel](https://www.kaggle.com/code/paddykb/tps-2022-10-fastai) - it is really awesome and you should take a look at it if you start working on this competition. \n\nWhen you are working on this kind of data, you will find some useful training data augmentations - shuffling players inside team, changing the teams in the mirror way (with changing targets as well) etc. But you can also do this for the test data - this is usually called Test Time Augmentation (TTA for short). You can do that because you know that the real target of the test data doesn't change if you do this kind of augmentation (all of them without mirror) so you can predict for the augmented test data several times and mean the received predictions.\n\nFor the neural networks, it's also will be better if you use multi-start technique - run the network training several times and also calculate the mean predictions for the runs.\n\nBoth of these techniques helped me to improve the @paddykb's kernel and now you can take a look of their implementation in my kernel (which is the cloned modified version of [this kernel](https://www.kaggle.com/code/paddykb/tps-2022-10-fastai): [link to kernel with TTA and multistart](https://www.kaggle.com/alexryzhkov/tps-2022-10-fastai-with-multistart-and-tta)\n\nHope this can help you to improve your scores.\n\nAlex",
      "votes": 14
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1983906": "Hi colleagues,\n\nFirst of all, I would like to mention the great appreciation for @paddykb and [his great kernel](https://www.kaggle.com/code/paddykb/tps-2022-10-fastai) - it is really awesome and you should take a look at it if you start working on this competition. \n\nWhen you are working on this kind of data, you will find some useful training data augmentations - shuffling players inside team, changing the teams in the mirror way (with changing targets as well) etc. But you can also do this for the test data - this is usually called Test Time Augmentation (TTA for short). You can do that because you know that the real target of the test data doesn't change if you do this kind of augmentation (all of them without mirror) so you can predict for the augmented test data several times and mean the received predictions.\n\nFor the neural networks, it's also will be better if you use multi-start technique - run the network training several times and also calculate the mean predictions for the runs.\n\nBoth of these techniques helped me to improve the @paddykb's kernel and now you can take a look of their implementation in my kernel (which is the cloned modified version of [this kernel](https://www.kaggle.com/code/paddykb/tps-2022-10-fastai): [link to kernel with TTA and multistart](https://www.kaggle.com/alexryzhkov/tps-2022-10-fastai-with-multistart-and-tta)\n\nHope this can help you to improve your scores.\n\nAlex"
  }
}