{
  "id": 332930,
  "title": "Tabular Data Augmentations",
  "url": "/competitions/amex-default-prediction/discussion/332930",
  "author_name": "",
  "post_date": "2022-06-24T01:55:47.734192100Z",
  "votes": 25,
  "comment_count": 4,
  "views": 0,
  "content": "<h3>Tabular Data Augmentations</h3>\n<p>We all know that augmentations work great on images (and sometimes on text &amp; time series). And it makes sense, a dog is still a dog even if it is rotated. And the network should understand this. <br>\nHowever, when we deal with tabular data the situation becomes a bit tricky: we can not \"rotate\" a table nor can we \"zoom\". <br>\nSo what can we do?</p>\n<h4>Simple Noise (jitter)</h4>\n<p>Simply put, we can simply add noise to the columns themselves.<br>\nOne simple improvement for this approach would be to take in consideration the std of the column itself when coming up with the noise we want to add. </p>\n<h4>Swap Noise</h4>\n<p>This approach was used in the past <a href=\"https://www.kaggle.com/code/springmanndaniel/1st-place-turn-your-data-into-daeta/report\" target=\"_blank\">multiple</a> <a href=\"https://www.kaggle.com/competitions/tabular-playground-series-apr-2021/discussion/235739\" target=\"_blank\">times</a> to win competitions. <br>\nThe trick is \"swapping\" cells of the same feature column with a different cell from the same feature column. (don't mix features)<br>\nThis makes it so the output would be a feature column that still contain values that make sense for the respective feature. </p>\n<p>Iv'e done a little bit of digging and I was able to find this function that had been used by the <a href=\"https://www.kaggle.com/code/jiangtt/tps-apr-2021-pseudo-labeling-voting-ensemble/notebook\" target=\"_blank\">winning</a> solution of TPS-APR 2021. </p>\n<pre><code>def apply_noise(df, p=.75):\n    should_not_swap = np.random.binomial(1, p, df.shape)\n    corrupted_df = df.where(should_not_swap == 1, np.random.permutation(df))\n    return corrupted_df\n</code></pre>\n<p><strong>Types of swap noise</strong></p>\n<p>If we consider the batch size and feature space be five (just for the animation example) and set the noise probaility to 20%, we get the following animations for each type.<br>\n[This means that 20% of the data is replaced by noise]</p>\n<p><strong>Columnwise swap noise</strong><br>\nThe idea of columwise noise is to noise only 20% of the batch by touching only full columns (check the animation below).</p>\n<p><img src=\"https://i.ibb.co/x5kwXKB/col-noise-small.gif\" alt=\"\"><br>\n<a href=\"https://www.kaggle.com/code/springmanndaniel/1st-place-turn-your-data-into-daeta/report\" target=\"_blank\">source</a></p>\n<p><strong>Randomized swap noise</strong><br>\nRandomized noise ads 20% noisy values in a random way (see animation below)</p>\n<p><img src=\"https://i.ibb.co/8M55Qd0/random-noise-small.gif\" alt=\"\"><br>\n<a href=\"https://www.kaggle.com/code/springmanndaniel/1st-place-turn-your-data-into-daeta/report\" target=\"_blank\">source</a></p>\n<p><strong>Rowwise swap noise</strong><br>\nAs the name indicates rowwise noise replaces 20% of data so that each and every row is beeing transformed by a certain amount of noise.</p>\n<p><img src=\"https://i.ibb.co/GPYYBFy/row-noise-small.gif\" alt=\"\"></p>\n<p><a href=\"https://www.kaggle.com/code/springmanndaniel/1st-place-turn-your-data-into-daeta/report\" target=\"_blank\">source</a></p>\n<h4>MixUp</h4>\n<p>Another type of augmentation we can easily use is coming streight from computer vision: <code>MixUp</code>.<br>\n<img src=\"https://i.ibb.co/xHpQGLM/1-Xqy-D5-OE47-Adqe-R6-Ke-Mg9-FQ.png\" alt=\"\"></p>\n<p>To do this, we simply create linear combinations of rows from the table (Don't forget to do this also for the target).</p>\n<h4>CutMix</h4>\n<p>And another simple augmentation we can use is again coming streight from computer vision: <code>MixUp</code>.<br>\n<img src=\"https://i.ibb.co/bBgx9wG/Selection-913.png\" alt=\"\"></p>\n<p>To do this, we simply create linear combinations of rows from the table (Don't forget to do this also for the target).</p>\n<h4>CTGan</h4>\n<p>This <a href=\"https://arxiv.org/abs/1907.00503\" target=\"_blank\">paper</a> proposed an architecture of a GAN for tabular data.<br>\nIt was used in the past on <a href=\"https://www.kaggle.com/gogo827jz/data-augmentation-ctgan\" target=\"_blank\">kaggle</a> to generate augmented data. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3177784%2F57ac10ed55722be3ec6250273932cf8f%2FCTGAN.png?generation=1604351202659684&amp;alt=media\" alt=\"\"></p>\n<p>Feel free to try it as well! </p>\n<p>Do you got any more ideas to add here?</p>",
  "messages": [
    {
      "id": "1831149",
      "postDate": "06/24/2022 01:55:47",
      "content": "<h3>Tabular Data Augmentations</h3>\n<p>We all know that augmentations work great on images (and sometimes on text &amp; time series). And it makes sense, a dog is still a dog even if it is rotated. And the network should understand this. <br>\nHowever, when we deal with tabular data the situation becomes a bit tricky: we can not \"rotate\" a table nor can we \"zoom\". <br>\nSo what can we do?</p>\n<h4>Simple Noise (jitter)</h4>\n<p>Simply put, we can simply add noise to the columns themselves.<br>\nOne simple improvement for this approach would be to take in consideration the std of the column itself when coming up with the noise we want to add. </p>\n<h4>Swap Noise</h4>\n<p>This approach was used in the past <a href=\"https://www.kaggle.com/code/springmanndaniel/1st-place-turn-your-data-into-daeta/report\" target=\"_blank\">multiple</a> <a href=\"https://www.kaggle.com/competitions/tabular-playground-series-apr-2021/discussion/235739\" target=\"_blank\">times</a> to win competitions. <br>\nThe trick is \"swapping\" cells of the same feature column with a different cell from the same feature column. (don't mix features)<br>\nThis makes it so the output would be a feature column that still contain values that make sense for the respective feature. </p>\n<p>Iv'e done a little bit of digging and I was able to find this function that had been used by the <a href=\"https://www.kaggle.com/code/jiangtt/tps-apr-2021-pseudo-labeling-voting-ensemble/notebook\" target=\"_blank\">winning</a> solution of TPS-APR 2021. </p>\n<pre><code>def apply_noise(df, p=.75):\n    should_not_swap = np.random.binomial(1, p, df.shape)\n    corrupted_df = df.where(should_not_swap == 1, np.random.permutation(df))\n    return corrupted_df\n</code></pre>\n<p><strong>Types of swap noise</strong></p>\n<p>If we consider the batch size and feature space be five (just for the animation example) and set the noise probaility to 20%, we get the following animations for each type.<br>\n[This means that 20% of the data is replaced by noise]</p>\n<p><strong>Columnwise swap noise</strong><br>\nThe idea of columwise noise is to noise only 20% of the batch by touching only full columns (check the animation below).</p>\n<p><img src=\"https://i.ibb.co/x5kwXKB/col-noise-small.gif\" alt=\"\"><br>\n<a href=\"https://www.kaggle.com/code/springmanndaniel/1st-place-turn-your-data-into-daeta/report\" target=\"_blank\">source</a></p>\n<p><strong>Randomized swap noise</strong><br>\nRandomized noise ads 20% noisy values in a random way (see animation below)</p>\n<p><img src=\"https://i.ibb.co/8M55Qd0/random-noise-small.gif\" alt=\"\"><br>\n<a href=\"https://www.kaggle.com/code/springmanndaniel/1st-place-turn-your-data-into-daeta/report\" target=\"_blank\">source</a></p>\n<p><strong>Rowwise swap noise</strong><br>\nAs the name indicates rowwise noise replaces 20% of data so that each and every row is beeing transformed by a certain amount of noise.</p>\n<p><img src=\"https://i.ibb.co/GPYYBFy/row-noise-small.gif\" alt=\"\"></p>\n<p><a href=\"https://www.kaggle.com/code/springmanndaniel/1st-place-turn-your-data-into-daeta/report\" target=\"_blank\">source</a></p>\n<h4>MixUp</h4>\n<p>Another type of augmentation we can easily use is coming streight from computer vision: <code>MixUp</code>.<br>\n<img src=\"https://i.ibb.co/xHpQGLM/1-Xqy-D5-OE47-Adqe-R6-Ke-Mg9-FQ.png\" alt=\"\"></p>\n<p>To do this, we simply create linear combinations of rows from the table (Don't forget to do this also for the target).</p>\n<h4>CutMix</h4>\n<p>And another simple augmentation we can use is again coming streight from computer vision: <code>MixUp</code>.<br>\n<img src=\"https://i.ibb.co/bBgx9wG/Selection-913.png\" alt=\"\"></p>\n<p>To do this, we simply create linear combinations of rows from the table (Don't forget to do this also for the target).</p>\n<h4>CTGan</h4>\n<p>This <a href=\"https://arxiv.org/abs/1907.00503\" target=\"_blank\">paper</a> proposed an architecture of a GAN for tabular data.<br>\nIt was used in the past on <a href=\"https://www.kaggle.com/gogo827jz/data-augmentation-ctgan\" target=\"_blank\">kaggle</a> to generate augmented data. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3177784%2F57ac10ed55722be3ec6250273932cf8f%2FCTGAN.png?generation=1604351202659684&amp;alt=media\" alt=\"\"></p>\n<p>Feel free to try it as well! </p>\n<p>Do you got any more ideas to add here?</p>",
      "rawMarkdown": "### Tabular Data Augmentations\n\nWe all know that augmentations work great on images (and sometimes on text & time series). And it makes sense, a dog is still a dog even if it is rotated. And the network should understand this. \nHowever, when we deal with tabular data the situation becomes a bit tricky: we can not \"rotate\" a table nor can we \"zoom\". \nSo what can we do?\n\n\n#### Simple Noise (jitter)\n\nSimply put, we can simply add noise to the columns themselves.\nOne simple improvement for this approach would be to take in consideration the std of the column itself when coming up with the noise we want to add. \n\n\n#### Swap Noise\n\nThis approach was used in the past [multiple](https://www.kaggle.com/code/springmanndaniel/1st-place-turn-your-data-into-daeta/report) [times](https://www.kaggle.com/competitions/tabular-playground-series-apr-2021/discussion/235739) to win competitions. \nThe trick is \"swapping\" cells of the same feature column with a different cell from the same feature column. (don't mix features)\nThis makes it so the output would be a feature column that still contain values that make sense for the respective feature. \n\n\nIv'e done a little bit of digging and I was able to find this function that had been used by the [winning](https://www.kaggle.com/code/jiangtt/tps-apr-2021-pseudo-labeling-voting-ensemble/notebook) solution of TPS-APR 2021. \n\n```python\ndef apply_noise(df, p=.75):\n    should_not_swap = np.random.binomial(1, p, df.shape)\n    corrupted_df = df.where(should_not_swap == 1, np.random.permutation(df))\n    return corrupted_df\n```\n\n**Types of swap noise**\n\nIf we consider the batch size and feature space be five (just for the animation example) and set the noise probaility to 20%, we get the following animations for each type.\n[This means that 20% of the data is replaced by noise]\n\n\n**Columnwise swap noise**\nThe idea of columwise noise is to noise only 20% of the batch by touching only full columns (check the animation below).\n\n![](https://i.ibb.co/x5kwXKB/col-noise-small.gif)\n[source](https://www.kaggle.com/code/springmanndaniel/1st-place-turn-your-data-into-daeta/report)\n\n\n**Randomized swap noise**\nRandomized noise ads 20% noisy values in a random way (see animation below)\n\n![](https://i.ibb.co/8M55Qd0/random-noise-small.gif)\n[source](https://www.kaggle.com/code/springmanndaniel/1st-place-turn-your-data-into-daeta/report)\n\n**Rowwise swap noise**\nAs the name indicates rowwise noise replaces 20% of data so that each and every row is beeing transformed by a certain amount of noise.\n\n![](https://i.ibb.co/GPYYBFy/row-noise-small.gif)\n\n[source](https://www.kaggle.com/code/springmanndaniel/1st-place-turn-your-data-into-daeta/report)\n\n#### MixUp\n\nAnother type of augmentation we can easily use is coming streight from computer vision: `MixUp`.\n![](https://i.ibb.co/xHpQGLM/1-Xqy-D5-OE47-Adqe-R6-Ke-Mg9-FQ.png)\n\nTo do this, we simply create linear combinations of rows from the table (Don't forget to do this also for the target).\n\n\n#### CutMix\n\nAnd another simple augmentation we can use is again coming streight from computer vision: `MixUp`.\n![](https://i.ibb.co/bBgx9wG/Selection-913.png)\n\nTo do this, we simply create linear combinations of rows from the table (Don't forget to do this also for the target).\n\n\n#### CTGan\n\nThis [paper](https://arxiv.org/abs/1907.00503) proposed an architecture of a GAN for tabular data.\nIt was used in the past on [kaggle](https://www.kaggle.com/gogo827jz/data-augmentation-ctgan) to generate augmented data. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3177784%2F57ac10ed55722be3ec6250273932cf8f%2FCTGAN.png?generation=1604351202659684&alt=media)\n\nFeel free to try it as well! \n\n\nDo you got any more ideas to add here?",
      "votes": null
    },
    {
      "id": "1831494",
      "postDate": "06/24/2022 07:49:17",
      "content": "<p>Devastating information once again! Thanks for sharing and more the great educational material you're providing in this competition</p>",
      "rawMarkdown": "Devastating information once again! Thanks for sharing and more the great educational material you're providing in this competition",
      "votes": null
    },
    {
      "id": "1831592",
      "postDate": "06/24/2022 09:37:39",
      "content": "<p>Nice Summary. I had not tryed any yet. in this competion. There are some many data. But I have worked with a deep tabular augmentation once from this git <a href=\"https://github.com/lschmiddey/deep_tabular_augmentation/tree/main/deep_tabular_augmentation\" target=\"_blank\">deep_tabular_augmentation model:</a></p>",
      "rawMarkdown": "Nice Summary. I had not tryed any yet. in this competion. There are some many data. But I have worked with a deep tabular augmentation once from this git [deep_tabular_augmentation model:](https://github.com/lschmiddey/deep_tabular_augmentation/tree/main/deep_tabular_augmentation)",
      "votes": null
    },
    {
      "id": "1832436",
      "postDate": "06/25/2022 03:46:44",
      "content": "<p>Thanks for bringing this up! Came to know about the augmentation technique for Tabular data for the first time!</p>",
      "rawMarkdown": "Thanks for bringing this up! Came to know about the augmentation technique for Tabular data for the first time!",
      "votes": null
    },
    {
      "id": "1834065",
      "postDate": "06/26/2022 15:36:54",
      "content": "<p>Thank you for sharing!</p>",
      "rawMarkdown": "Thank you for sharing!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1831494,
      "author_name": "nymfree",
      "author_url": "",
      "post_date": "06/24/2022 07:49:17",
      "content": "<p>Devastating information once again! Thanks for sharing and more the great educational material you're providing in this competition</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1831592,
      "author_name": "paulojunqueira",
      "author_url": "",
      "post_date": "06/24/2022 09:37:39",
      "content": "<p>Nice Summary. I had not tryed any yet. in this competion. There are some many data. But I have worked with a deep tabular augmentation once from this git <a href=\"https://github.com/lschmiddey/deep_tabular_augmentation/tree/main/deep_tabular_augmentation\" target=\"_blank\">deep_tabular_augmentation model:</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1832436,
      "author_name": "towhidultonmoy",
      "author_url": "",
      "post_date": "06/25/2022 03:46:44",
      "content": "<p>Thanks for bringing this up! Came to know about the augmentation technique for Tabular data for the first time!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1834065,
      "author_name": "jinkyh",
      "author_url": "",
      "post_date": "06/26/2022 15:36:54",
      "content": "<p>Thank you for sharing!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1831149": "### Tabular Data Augmentations\n\nWe all know that augmentations work great on images (and sometimes on text & time series). And it makes sense, a dog is still a dog even if it is rotated. And the network should understand this. \nHowever, when we deal with tabular data the situation becomes a bit tricky: we can not \"rotate\" a table nor can we \"zoom\". \nSo what can we do?\n\n\n#### Simple Noise (jitter)\n\nSimply put, we can simply add noise to the columns themselves.\nOne simple improvement for this approach would be to take in consideration the std of the column itself when coming up with the noise we want to add. \n\n\n#### Swap Noise\n\nThis approach was used in the past [multiple](https://www.kaggle.com/code/springmanndaniel/1st-place-turn-your-data-into-daeta/report) [times](https://www.kaggle.com/competitions/tabular-playground-series-apr-2021/discussion/235739) to win competitions. \nThe trick is \"swapping\" cells of the same feature column with a different cell from the same feature column. (don't mix features)\nThis makes it so the output would be a feature column that still contain values that make sense for the respective feature. \n\n\nIv'e done a little bit of digging and I was able to find this function that had been used by the [winning](https://www.kaggle.com/code/jiangtt/tps-apr-2021-pseudo-labeling-voting-ensemble/notebook) solution of TPS-APR 2021. \n\n```python\ndef apply_noise(df, p=.75):\n    should_not_swap = np.random.binomial(1, p, df.shape)\n    corrupted_df = df.where(should_not_swap == 1, np.random.permutation(df))\n    return corrupted_df\n```\n\n**Types of swap noise**\n\nIf we consider the batch size and feature space be five (just for the animation example) and set the noise probaility to 20%, we get the following animations for each type.\n[This means that 20% of the data is replaced by noise]\n\n\n**Columnwise swap noise**\nThe idea of columwise noise is to noise only 20% of the batch by touching only full columns (check the animation below).\n\n![](https://i.ibb.co/x5kwXKB/col-noise-small.gif)\n[source](https://www.kaggle.com/code/springmanndaniel/1st-place-turn-your-data-into-daeta/report)\n\n\n**Randomized swap noise**\nRandomized noise ads 20% noisy values in a random way (see animation below)\n\n![](https://i.ibb.co/8M55Qd0/random-noise-small.gif)\n[source](https://www.kaggle.com/code/springmanndaniel/1st-place-turn-your-data-into-daeta/report)\n\n**Rowwise swap noise**\nAs the name indicates rowwise noise replaces 20% of data so that each and every row is beeing transformed by a certain amount of noise.\n\n![](https://i.ibb.co/GPYYBFy/row-noise-small.gif)\n\n[source](https://www.kaggle.com/code/springmanndaniel/1st-place-turn-your-data-into-daeta/report)\n\n#### MixUp\n\nAnother type of augmentation we can easily use is coming streight from computer vision: `MixUp`.\n![](https://i.ibb.co/xHpQGLM/1-Xqy-D5-OE47-Adqe-R6-Ke-Mg9-FQ.png)\n\nTo do this, we simply create linear combinations of rows from the table (Don't forget to do this also for the target).\n\n\n#### CutMix\n\nAnd another simple augmentation we can use is again coming streight from computer vision: `MixUp`.\n![](https://i.ibb.co/bBgx9wG/Selection-913.png)\n\nTo do this, we simply create linear combinations of rows from the table (Don't forget to do this also for the target).\n\n\n#### CTGan\n\nThis [paper](https://arxiv.org/abs/1907.00503) proposed an architecture of a GAN for tabular data.\nIt was used in the past on [kaggle](https://www.kaggle.com/gogo827jz/data-augmentation-ctgan) to generate augmented data. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3177784%2F57ac10ed55722be3ec6250273932cf8f%2FCTGAN.png?generation=1604351202659684&alt=media)\n\nFeel free to try it as well! \n\n\nDo you got any more ideas to add here?",
    "1831494": "Devastating information once again! Thanks for sharing and more the great educational material you're providing in this competition",
    "1831592": "Nice Summary. I had not tryed any yet. in this competion. There are some many data. But I have worked with a deep tabular augmentation once from this git [deep_tabular_augmentation model:](https://github.com/lschmiddey/deep_tabular_augmentation/tree/main/deep_tabular_augmentation)",
    "1832436": "Thanks for bringing this up! Came to know about the augmentation technique for Tabular data for the first time!",
    "1834065": "Thank you for sharing!"
  },
  "source": "meta"
}