{
  "id": 246163,
  "title": "[Ablation Study] What's the best Alpha value for Mixup Augmentation?",
  "url": "/competitions/seti-breakthrough-listen/discussion/246163",
  "author_name": "Ayush Thakur",
  "post_date": "2021-06-14T09:10:06.750000",
  "votes": 21,
  "comment_count": 6,
  "views": 0,
  "content": "<h1>Introduction</h1>\n<p>The mixup augmentation mixes two images pixel-wise and mixes their labels as well. This is done by weighted element-wise sum where the weight is sampled from the <a href=\"https://en.wikipedia.org/wiki/Beta_distribution\" target=\"_blank\">Beta Distribution</a>. </p>\n<p>The Beta distribution depends on two parameters - <code>alpha</code> and <code>beta</code>. In the context of Mixup, the <code>alpha</code> and <code>beta</code> take the same value and the value is less or equal to 1.0. You can play with this interactive chart <a href=\"https://keisan.casio.com/exec/system/1180573226\" target=\"_blank\">here</a>.</p>\n<p>In this experiment we will use different values of alpha (beta) and find out:</p>\n<ul>\n<li>if there is any effect of alpha on this dataset,</li>\n<li>if yes, what's the optimal value to use. </li>\n</ul>\n<h1>Experiment Setup</h1>\n<ul>\n<li>Framework: TensorFlow</li>\n<li>Data: 5000 examples randomly shuffled. The resulting data distribution is close to the full training data distribution.</li>\n<li>Backbone architecture: EfficientNetB0</li>\n<li>Experiment Tracking: I used [Weights and Biases]() for tracking all the experiments.</li>\n<li>Trained for: Each experiment was run 3 times to get the mean and standard deviation.</li>\n<li>Regularization: Trained with early stopping with the patience of 5 epochs.</li>\n<li>Other Configs: Adam optimizer was used with a learning rate of 1e-3. </li>\n</ul>\n<h1>Alpha vs AUC-ROC</h1>\n<ul>\n<li>Looking at the chart below,** every value of alpha is giving \"almost\" the same validation AUC-ROC score**. </li>\n</ul>\n<p><img src=\"https://i.imgur.com/1ylEDyO.png\" alt=\"img\"><br>\n(<a href=\"http://wandb.me/kaggle-seti-alpha-mixup\" target=\"_blank\">Source</a>)</p>\n<ul>\n<li>By sorting, the AUC-ROC value on the W&amp;B dashboard, we see that the <strong>alpha value of 0.2 is giving a slightly better score</strong>.  This is followed by 0.4, 1.0, 0.6, and finally 0.8. </li>\n</ul>\n<p><img src=\"https://i.imgur.com/uNeu2Os.png\" alt=\"img\"><br>\n(<a href=\"http://wandb.me/kaggle-seti-alpha-mixup\" target=\"_blank\">Source</a>)</p>\n<h1>Conclusion</h1>\n<ul>\n<li>There is no significant effect of the value of alpha with this dataset. </li>\n<li>Either use 0.2 or 1.0 as the value of alpha for Mixup augmentation. </li>\n</ul>\n<h3><a href=\"http://wandb.me/kernel-seti-exp\" target=\"_blank\">Kernel here</a> | <a href=\"http://wandb.me/kaggle-seti-alpha-mixup\" target=\"_blank\">W&amp;B Dashboad here</a></h3>",
  "messages": [
    {
      "id": 1348791,
      "postDate": "2021-06-14T09:10:06.750Z",
      "content": "<h1>Introduction</h1>\n<p>The mixup augmentation mixes two images pixel-wise and mixes their labels as well. This is done by weighted element-wise sum where the weight is sampled from the <a href=\"https://en.wikipedia.org/wiki/Beta_distribution\" target=\"_blank\">Beta Distribution</a>. </p>\n<p>The Beta distribution depends on two parameters - <code>alpha</code> and <code>beta</code>. In the context of Mixup, the <code>alpha</code> and <code>beta</code> take the same value and the value is less or equal to 1.0. You can play with this interactive chart <a href=\"https://keisan.casio.com/exec/system/1180573226\" target=\"_blank\">here</a>.</p>\n<p>In this experiment we will use different values of alpha (beta) and find out:</p>\n<ul>\n<li>if there is any effect of alpha on this dataset,</li>\n<li>if yes, what's the optimal value to use. </li>\n</ul>\n<h1>Experiment Setup</h1>\n<ul>\n<li>Framework: TensorFlow</li>\n<li>Data: 5000 examples randomly shuffled. The resulting data distribution is close to the full training data distribution.</li>\n<li>Backbone architecture: EfficientNetB0</li>\n<li>Experiment Tracking: I used [Weights and Biases]() for tracking all the experiments.</li>\n<li>Trained for: Each experiment was run 3 times to get the mean and standard deviation.</li>\n<li>Regularization: Trained with early stopping with the patience of 5 epochs.</li>\n<li>Other Configs: Adam optimizer was used with a learning rate of 1e-3. </li>\n</ul>\n<h1>Alpha vs AUC-ROC</h1>\n<ul>\n<li>Looking at the chart below,** every value of alpha is giving \"almost\" the same validation AUC-ROC score**. </li>\n</ul>\n<p><img src=\"https://i.imgur.com/1ylEDyO.png\" alt=\"img\"><br>\n(<a href=\"http://wandb.me/kaggle-seti-alpha-mixup\" target=\"_blank\">Source</a>)</p>\n<ul>\n<li>By sorting, the AUC-ROC value on the W&amp;B dashboard, we see that the <strong>alpha value of 0.2 is giving a slightly better score</strong>.  This is followed by 0.4, 1.0, 0.6, and finally 0.8. </li>\n</ul>\n<p><img src=\"https://i.imgur.com/uNeu2Os.png\" alt=\"img\"><br>\n(<a href=\"http://wandb.me/kaggle-seti-alpha-mixup\" target=\"_blank\">Source</a>)</p>\n<h1>Conclusion</h1>\n<ul>\n<li>There is no significant effect of the value of alpha with this dataset. </li>\n<li>Either use 0.2 or 1.0 as the value of alpha for Mixup augmentation. </li>\n</ul>\n<h3><a href=\"http://wandb.me/kernel-seti-exp\" target=\"_blank\">Kernel here</a> | <a href=\"http://wandb.me/kaggle-seti-alpha-mixup\" target=\"_blank\">W&amp;B Dashboad here</a></h3>",
      "rawMarkdown": "# Introduction\n\nThe mixup augmentation mixes two images pixel-wise and mixes their labels as well. This is done by weighted element-wise sum where the weight is sampled from the [Beta Distribution](https://en.wikipedia.org/wiki/Beta_distribution). \n\nThe Beta distribution depends on two parameters - `alpha` and `beta`. In the context of Mixup, the `alpha` and `beta` take the same value and the value is less or equal to 1.0. You can play with this interactive chart [here](https://keisan.casio.com/exec/system/1180573226).\n\nIn this experiment we will use different values of alpha (beta) and find out:\n* if there is any effect of alpha on this dataset,\n* if yes, what's the optimal value to use. \n\n# Experiment Setup\n\n* Framework: TensorFlow\n* Data: 5000 examples randomly shuffled. The resulting data distribution is close to the full training data distribution.\n* Backbone architecture: EfficientNetB0\n* Experiment Tracking: I used [Weights and Biases]() for tracking all the experiments.\n* Trained for: Each experiment was run 3 times to get the mean and standard deviation.\n* Regularization: Trained with early stopping with the patience of 5 epochs.\n* Other Configs: Adam optimizer was used with a learning rate of 1e-3. \n\n# Alpha vs AUC-ROC\n\n* Looking at the chart below,** every value of alpha is giving \"almost\" the same validation AUC-ROC score**. \n\n![img](https://i.imgur.com/1ylEDyO.png)\n([Source](http://wandb.me/kaggle-seti-alpha-mixup))\n\n* By sorting, the AUC-ROC value on the W&B dashboard, we see that the **alpha value of 0.2 is giving a slightly better score**.  This is followed by 0.4, 1.0, 0.6, and finally 0.8. \n\n![img](https://i.imgur.com/uNeu2Os.png)\n([Source](http://wandb.me/kaggle-seti-alpha-mixup))\n\n# Conclusion\n\n* There is no significant effect of the value of alpha with this dataset. \n* Either use 0.2 or 1.0 as the value of alpha for Mixup augmentation. \n\n### [Kernel here](http://wandb.me/kernel-seti-exp) | [W&B Dashboad here](http://wandb.me/kaggle-seti-alpha-mixup) ",
      "votes": 21
    },
    {
      "id": 1361117,
      "postDate": "2021-06-22T15:04:55.770Z",
      "content": "<p>This is great work, <a href=\"https://www.kaggle.com/ayuraj\" target=\"_blank\">@ayuraj</a>! I'm sorry I don't really understand what is the benefit of using wandb dashboard. Can you please tell me how has it benefitted you and how is it different from TensorBoard? Thanks :)</p>",
      "rawMarkdown": "This is great work, @ayuraj! I'm sorry I don't really understand what is the benefit of using wandb dashboard. Can you please tell me how has it benefitted you and how is it different from TensorBoard? Thanks :)",
      "votes": 1,
      "replies": [
        {
          "id": 1361790,
          "postDate": "2021-06-23T04:16:38.600Z",
          "content": "<p>Hey, <a href=\"https://www.kaggle.com/yashraizada\" target=\"_blank\">@yashraizada</a> thanks. </p>\n<p>W&amp;B dashboard is like a GitHub repository of all the experiments I conduct. Experiments can be preparing datasets, training models, doing ablation studies, hyperparameter optimizations, etc. With everything in one place, it's easier to compare one experiment from another, get insights that we usually can't get by maintaining experiments on a spreadsheet (this is manual extensive). The best part about this dashboard in my opinion is the ease of sharing this with teammates, friends, and colleagues. </p>\n<p>Now how is it different from TFboard?</p>\n<ul>\n<li>There is no issue of ports while logging runs. I used to use TFboard 2 yrs back and I had to set it up properly on different platforms differently, there were port issues, etc. </li>\n<li>W&amp;B is logging everything on a server making it platform agnostic. I can run the same code instrumented with W&amp;B on colab, kaggle, or GCP instance without any change in the W&amp;B code. </li>\n<li>I can write short reports in the dashboard itself. This way I keep a log of every new insight I get by running a bunch of experiments. The best part is that I can share this report with anyone. You can think this report to be a blog/forum post distilled with key informations. You can find one such report in this <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/245152\" target=\"_blank\">discussion post</a>. </li>\n<li>W&amp;B enables dataset and model version control using W&amp;B Artifacts. I can access this from the dashboard itself. I find it really useful when training a bunch of models. The image shows a dataset version control using W&amp;B artifacts that I used for Human Protein Atlas competition. <br>\n<img src=\"https://i.imgur.com/HwhhSip.png\" alt=\"Imgur\"></li>\n</ul>\n<p><a href=\"https://docs.wandb.ai/guides/integrations/tensorboard#how-is-w-and-b-different-from-tensorboard\" target=\"_blank\">You can get more information about the differences in this post.</a></p>",
          "rawMarkdown": "Hey, @yashraizada thanks. \n\nW&B dashboard is like a GitHub repository of all the experiments I conduct. Experiments can be preparing datasets, training models, doing ablation studies, hyperparameter optimizations, etc. With everything in one place, it's easier to compare one experiment from another, get insights that we usually can't get by maintaining experiments on a spreadsheet (this is manual extensive). The best part about this dashboard in my opinion is the ease of sharing this with teammates, friends, and colleagues. \n\nNow how is it different from TFboard?\n* There is no issue of ports while logging runs. I used to use TFboard 2 yrs back and I had to set it up properly on different platforms differently, there were port issues, etc. \n* W&B is logging everything on a server making it platform agnostic. I can run the same code instrumented with W&B on colab, kaggle, or GCP instance without any change in the W&B code. \n* I can write short reports in the dashboard itself. This way I keep a log of every new insight I get by running a bunch of experiments. The best part is that I can share this report with anyone. You can think this report to be a blog/forum post distilled with key informations. You can find one such report in this [discussion post](https://www.kaggle.com/c/seti-breakthrough-listen/discussion/245152). \n* W&B enables dataset and model version control using W&B Artifacts. I can access this from the dashboard itself. I find it really useful when training a bunch of models. The image shows a dataset version control using W&B artifacts that I used for Human Protein Atlas competition. \n![Imgur](https://i.imgur.com/HwhhSip.png)\n\n[You can get more information about the differences in this post.](https://docs.wandb.ai/guides/integrations/tensorboard#how-is-w-and-b-different-from-tensorboard)",
          "votes": 1
        }
      ]
    },
    {
      "id": 1349035,
      "postDate": "2021-06-14T13:28:49.300Z",
      "content": "<p>Insightful as always. Thanks for the experiments. </p>",
      "rawMarkdown": "Insightful as always. Thanks for the experiments. ",
      "votes": 1,
      "replies": [
        {
          "id": 1349790,
          "postDate": "2021-06-15T04:56:11.957Z",
          "content": "<p>Your welcome Hongnan. :)</p>",
          "rawMarkdown": "Your welcome Hongnan. :)"
        }
      ]
    },
    {
      "id": 1496444,
      "postDate": "2021-08-30T11:40:45.333Z",
      "content": "<p>Thanks for sharing. I am sorry I have some questions, but can you help me answer the following doubt?</p>\n<p>As you mentioned:<code>Data: 5000 examples randomly shuffled. The resulting data distribution is close to the full training data distribution.</code> But how can we ensure that the sampled data and the complete training data are the same distribution. </p>\n<p>I have tried <code>train = pd.read_csv('../data/train_labels.csv').sample(10000).reset_index(drop=True)</code>, but found some tricks worked in the sampled data didn't workd in the whole training data. I'm not sure if it happened because of their distribution is different, and if their distribution is differnent because of sampling like above.</p>",
      "rawMarkdown": "Thanks for sharing. I am sorry I have some questions, but can you help me answer the following doubt?\n\nAs you mentioned:`Data: 5000 examples randomly shuffled. The resulting data distribution is close to the full training data distribution.` But how can we ensure that the sampled data and the complete training data are the same distribution. \n\nI have tried `train = pd.read_csv('../data/train_labels.csv').sample(10000).reset_index(drop=True) `, but found some tricks worked in the sampled data didn't workd in the whole training data. I'm not sure if it happened because of their distribution is different, and if their distribution is differnent because of sampling like above."
    },
    {
      "id": 1350051,
      "postDate": "2021-06-15T08:34:51.537Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1361117,
      "author_name": "Yash Raizada",
      "author_url": "",
      "post_date": "2021-06-22T15:04:55.770000",
      "content": "<p>This is great work, <a href=\"https://www.kaggle.com/ayuraj\" target=\"_blank\">@ayuraj</a>! I'm sorry I don't really understand what is the benefit of using wandb dashboard. Can you please tell me how has it benefitted you and how is it different from TensorBoard? Thanks :)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1361790,
          "author_name": "Ayush Thakur",
          "author_url": "",
          "post_date": "2021-06-23T04:16:38.600000",
          "content": "<p>Hey, <a href=\"https://www.kaggle.com/yashraizada\" target=\"_blank\">@yashraizada</a> thanks. </p>\n<p>W&amp;B dashboard is like a GitHub repository of all the experiments I conduct. Experiments can be preparing datasets, training models, doing ablation studies, hyperparameter optimizations, etc. With everything in one place, it's easier to compare one experiment from another, get insights that we usually can't get by maintaining experiments on a spreadsheet (this is manual extensive). The best part about this dashboard in my opinion is the ease of sharing this with teammates, friends, and colleagues. </p>\n<p>Now how is it different from TFboard?</p>\n<ul>\n<li>There is no issue of ports while logging runs. I used to use TFboard 2 yrs back and I had to set it up properly on different platforms differently, there were port issues, etc. </li>\n<li>W&amp;B is logging everything on a server making it platform agnostic. I can run the same code instrumented with W&amp;B on colab, kaggle, or GCP instance without any change in the W&amp;B code. </li>\n<li>I can write short reports in the dashboard itself. This way I keep a log of every new insight I get by running a bunch of experiments. The best part is that I can share this report with anyone. You can think this report to be a blog/forum post distilled with key informations. You can find one such report in this <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/245152\" target=\"_blank\">discussion post</a>. </li>\n<li>W&amp;B enables dataset and model version control using W&amp;B Artifacts. I can access this from the dashboard itself. I find it really useful when training a bunch of models. The image shows a dataset version control using W&amp;B artifacts that I used for Human Protein Atlas competition. <br>\n<img src=\"https://i.imgur.com/HwhhSip.png\" alt=\"Imgur\"></li>\n</ul>\n<p><a href=\"https://docs.wandb.ai/guides/integrations/tensorboard#how-is-w-and-b-different-from-tensorboard\" target=\"_blank\">You can get more information about the differences in this post.</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1349035,
      "author_name": "gao-hongnan",
      "author_url": "",
      "post_date": "2021-06-14T13:28:49.300000",
      "content": "<p>Insightful as always. Thanks for the experiments. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1349790,
          "author_name": "Ayush Thakur",
          "author_url": "",
          "post_date": "2021-06-15T04:56:11.957000",
          "content": "<p>Your welcome Hongnan. :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1496444,
      "author_name": "Gainover",
      "author_url": "",
      "post_date": "2021-08-30T11:40:45.333000",
      "content": "<p>Thanks for sharing. I am sorry I have some questions, but can you help me answer the following doubt?</p>\n<p>As you mentioned:<code>Data: 5000 examples randomly shuffled. The resulting data distribution is close to the full training data distribution.</code> But how can we ensure that the sampled data and the complete training data are the same distribution. </p>\n<p>I have tried <code>train = pd.read_csv('../data/train_labels.csv').sample(10000).reset_index(drop=True)</code>, but found some tricks worked in the sampled data didn't workd in the whole training data. I'm not sure if it happened because of their distribution is different, and if their distribution is differnent because of sampling like above.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1350051,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-06-15T08:34:51.537000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1348791": "# Introduction\n\nThe mixup augmentation mixes two images pixel-wise and mixes their labels as well. This is done by weighted element-wise sum where the weight is sampled from the [Beta Distribution](https://en.wikipedia.org/wiki/Beta_distribution). \n\nThe Beta distribution depends on two parameters - `alpha` and `beta`. In the context of Mixup, the `alpha` and `beta` take the same value and the value is less or equal to 1.0. You can play with this interactive chart [here](https://keisan.casio.com/exec/system/1180573226).\n\nIn this experiment we will use different values of alpha (beta) and find out:\n* if there is any effect of alpha on this dataset,\n* if yes, what's the optimal value to use. \n\n# Experiment Setup\n\n* Framework: TensorFlow\n* Data: 5000 examples randomly shuffled. The resulting data distribution is close to the full training data distribution.\n* Backbone architecture: EfficientNetB0\n* Experiment Tracking: I used [Weights and Biases]() for tracking all the experiments.\n* Trained for: Each experiment was run 3 times to get the mean and standard deviation.\n* Regularization: Trained with early stopping with the patience of 5 epochs.\n* Other Configs: Adam optimizer was used with a learning rate of 1e-3. \n\n# Alpha vs AUC-ROC\n\n* Looking at the chart below,** every value of alpha is giving \"almost\" the same validation AUC-ROC score**. \n\n![img](https://i.imgur.com/1ylEDyO.png)\n([Source](http://wandb.me/kaggle-seti-alpha-mixup))\n\n* By sorting, the AUC-ROC value on the W&B dashboard, we see that the **alpha value of 0.2 is giving a slightly better score**.  This is followed by 0.4, 1.0, 0.6, and finally 0.8. \n\n![img](https://i.imgur.com/uNeu2Os.png)\n([Source](http://wandb.me/kaggle-seti-alpha-mixup))\n\n# Conclusion\n\n* There is no significant effect of the value of alpha with this dataset. \n* Either use 0.2 or 1.0 as the value of alpha for Mixup augmentation. \n\n### [Kernel here](http://wandb.me/kernel-seti-exp) | [W&B Dashboad here](http://wandb.me/kaggle-seti-alpha-mixup) ",
    "1361117": "This is great work, @ayuraj! I'm sorry I don't really understand what is the benefit of using wandb dashboard. Can you please tell me how has it benefitted you and how is it different from TensorBoard? Thanks :)",
    "1349035": "Insightful as always. Thanks for the experiments. ",
    "1496444": "Thanks for sharing. I am sorry I have some questions, but can you help me answer the following doubt?\n\nAs you mentioned:`Data: 5000 examples randomly shuffled. The resulting data distribution is close to the full training data distribution.` But how can we ensure that the sampled data and the complete training data are the same distribution. \n\nI have tried `train = pd.read_csv('../data/train_labels.csv').sample(10000).reset_index(drop=True) `, but found some tricks worked in the sampled data didn't workd in the whole training data. I'm not sure if it happened because of their distribution is different, and if their distribution is differnent because of sampling like above.",
    "1350051": ""
  }
}