{
  "id": 549712,
  "title": "Test Time Augmentation",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/549712",
  "author_name": "",
  "post_date": "2024-12-03T15:21:39.311511200Z",
  "votes": 4,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I have not seen any mention of test time augmentation (TTA), but it can be helpful here, depending on how you trained your models. But you need to be careful. Here are some of my ideas I am working through. I would share notebooks, but the size of the data makes this challenging so a discussion post must suffice.</p>\n<p><strong>What is it?</strong><br>\n TTA is a technique that involves applying various transformations (augmentations) to the input data at test time, making predictions for each augmented version, then aggregating these predictions to produce a final result. It is common in the computer vision competitions I have seen on Kaggle, because augmentations are heavily applied during training to artificially increase sample size and improve generalization. </p>\n<p><strong>When is it helpful?</strong><br>\n1) When test data may differ slightly from training data (distribution shift). 2) When train time augmentations were used. E.g. if you applied Gaussian noise in training, it is a candidate for TTA. If you applied random distribution shifts during training, it is a candidate for TTA. You get the idea. </p>\n<p><strong>Guidelines</strong><br>\nEvaluate augmentations and <code>num_augmentations</code> during cross validation to see if they are actually helping. Furthermore, weigh the performance increase against the increased inference time, which is a limiting factor for code competitions like this one. </p>\n<p><strong>Simple code example</strong></p>\n<pre><code> ():\n    predictions = []\n\n     _  (num_augmentations):\n        \n        noisy_X_test = X_test + np.random.normal(, noise_stddev, X_test.shape)\n        preds = model.predict(noisy_X_test, batch_size=batch_size)\n        predictions.append(preds)\n     np.mean(predictions, axis=)\n</code></pre>",
  "messages": [
    {
      "id": "3062427",
      "postDate": "12/03/2024 15:21:39",
      "content": "<p>I have not seen any mention of test time augmentation (TTA), but it can be helpful here, depending on how you trained your models. But you need to be careful. Here are some of my ideas I am working through. I would share notebooks, but the size of the data makes this challenging so a discussion post must suffice.</p>\n<p><strong>What is it?</strong><br>\n TTA is a technique that involves applying various transformations (augmentations) to the input data at test time, making predictions for each augmented version, then aggregating these predictions to produce a final result. It is common in the computer vision competitions I have seen on Kaggle, because augmentations are heavily applied during training to artificially increase sample size and improve generalization. </p>\n<p><strong>When is it helpful?</strong><br>\n1) When test data may differ slightly from training data (distribution shift). 2) When train time augmentations were used. E.g. if you applied Gaussian noise in training, it is a candidate for TTA. If you applied random distribution shifts during training, it is a candidate for TTA. You get the idea. </p>\n<p><strong>Guidelines</strong><br>\nEvaluate augmentations and <code>num_augmentations</code> during cross validation to see if they are actually helping. Furthermore, weigh the performance increase against the increased inference time, which is a limiting factor for code competitions like this one. </p>\n<p><strong>Simple code example</strong></p>\n<pre><code> ():\n    predictions = []\n\n     _  (num_augmentations):\n        \n        noisy_X_test = X_test + np.random.normal(, noise_stddev, X_test.shape)\n        preds = model.predict(noisy_X_test, batch_size=batch_size)\n        predictions.append(preds)\n     np.mean(predictions, axis=)\n</code></pre>",
      "rawMarkdown": "I have not seen any mention of test time augmentation (TTA), but it can be helpful here, depending on how you trained your models. But you need to be careful. Here are some of my ideas I am working through. I would share notebooks, but the size of the data makes this challenging so a discussion post must suffice.\n\n**What is it?**\n TTA is a technique that involves applying various transformations (augmentations) to the input data at test time, making predictions for each augmented version, then aggregating these predictions to produce a final result. It is common in the computer vision competitions I have seen on Kaggle, because augmentations are heavily applied during training to artificially increase sample size and improve generalization. \n\n**When is it helpful?**\n1) When test data may differ slightly from training data (distribution shift). 2) When train time augmentations were used. E.g. if you applied Gaussian noise in training, it is a candidate for TTA. If you applied random distribution shifts during training, it is a candidate for TTA. You get the idea. \n\n**Guidelines**\nEvaluate augmentations and `num_augmentations` during cross validation to see if they are actually helping. Furthermore, weigh the performance increase against the increased inference time, which is a limiting factor for code competitions like this one. \n\n**Simple code example**\n```python\ndef gaussian_noise_tta(model, X_test, num_augmentations, noise_stddev, batch_size):\n    predictions = []\n\n    for _ in range(num_augmentations):\n        # Add Gaussian noise\n        noisy_X_test = X_test + np.random.normal(0, noise_stddev, X_test.shape)\n        preds = model.predict(noisy_X_test, batch_size=batch_size)\n        predictions.append(preds)\n    return np.mean(predictions, axis=0)\n```",
      "votes": null
    },
    {
      "id": "3062492",
      "postDate": "12/03/2024 16:08:14",
      "content": "<p>wonderful ideas</p>",
      "rawMarkdown": "wonderful ideas",
      "votes": null
    },
    {
      "id": "3062507",
      "postDate": "12/03/2024 16:17:42",
      "content": "<p>Do you have some guidelines on selecting standard deviation for Gaussian noise augmentation for this type of data?</p>",
      "rawMarkdown": "Do you have some guidelines on selecting standard deviation for Gaussian noise augmentation for this type of data?",
      "votes": null
    },
    {
      "id": "3062574",
      "postDate": "12/03/2024 17:20:24",
      "content": "<p>It should be roughly the same as what was applied during training. If you look at Yirun's winning write up from last competition, he used <code>0.035</code>. I would recommend trying 0 first, then increasing by <code>0.01</code> increments to find the sweet spot, then use that value in TTA. </p>\n<p>Also, make sure the testing data is scaled the same as the training data when applying noise TTA. You want to replicate training augmentation as much as possible. </p>",
      "rawMarkdown": "It should be roughly the same as what was applied during training. If you look at Yirun's winning write up from last competition, he used `0.035`. I would recommend trying 0 first, then increasing by `0.01` increments to find the sweet spot, then use that value in TTA. \n\nAlso, make sure the testing data is scaled the same as the training data when applying noise TTA. You want to replicate training augmentation as much as possible.",
      "votes": null
    },
    {
      "id": "3063988",
      "postDate": "12/05/2024 05:00:50",
      "content": "<p>Maybe… Let’s what happen going on!</p>",
      "rawMarkdown": "Maybe… Let’s what happen going on!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3062492,
      "author_name": "aichangeworld",
      "author_url": "",
      "post_date": "12/03/2024 16:08:14",
      "content": "<p>wonderful ideas</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3062507,
      "author_name": "redfoongus",
      "author_url": "",
      "post_date": "12/03/2024 16:17:42",
      "content": "<p>Do you have some guidelines on selecting standard deviation for Gaussian noise augmentation for this type of data?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3062574,
          "author_name": "tuckerarrants",
          "author_url": "",
          "post_date": "12/03/2024 17:20:24",
          "content": "<p>It should be roughly the same as what was applied during training. If you look at Yirun's winning write up from last competition, he used <code>0.035</code>. I would recommend trying 0 first, then increasing by <code>0.01</code> increments to find the sweet spot, then use that value in TTA. </p>\n<p>Also, make sure the testing data is scaled the same as the training data when applying noise TTA. You want to replicate training augmentation as much as possible. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3063988,
      "author_name": "junhanzangai",
      "author_url": "",
      "post_date": "12/05/2024 05:00:50",
      "content": "<p>Maybe… Let’s what happen going on!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3062427": "I have not seen any mention of test time augmentation (TTA), but it can be helpful here, depending on how you trained your models. But you need to be careful. Here are some of my ideas I am working through. I would share notebooks, but the size of the data makes this challenging so a discussion post must suffice.\n\n**What is it?**\n TTA is a technique that involves applying various transformations (augmentations) to the input data at test time, making predictions for each augmented version, then aggregating these predictions to produce a final result. It is common in the computer vision competitions I have seen on Kaggle, because augmentations are heavily applied during training to artificially increase sample size and improve generalization. \n\n**When is it helpful?**\n1) When test data may differ slightly from training data (distribution shift). 2) When train time augmentations were used. E.g. if you applied Gaussian noise in training, it is a candidate for TTA. If you applied random distribution shifts during training, it is a candidate for TTA. You get the idea. \n\n**Guidelines**\nEvaluate augmentations and `num_augmentations` during cross validation to see if they are actually helping. Furthermore, weigh the performance increase against the increased inference time, which is a limiting factor for code competitions like this one. \n\n**Simple code example**\n```python\ndef gaussian_noise_tta(model, X_test, num_augmentations, noise_stddev, batch_size):\n    predictions = []\n\n    for _ in range(num_augmentations):\n        # Add Gaussian noise\n        noisy_X_test = X_test + np.random.normal(0, noise_stddev, X_test.shape)\n        preds = model.predict(noisy_X_test, batch_size=batch_size)\n        predictions.append(preds)\n    return np.mean(predictions, axis=0)\n```",
    "3062492": "wonderful ideas",
    "3062507": "Do you have some guidelines on selecting standard deviation for Gaussian noise augmentation for this type of data?",
    "3062574": "It should be roughly the same as what was applied during training. If you look at Yirun's winning write up from last competition, he used `0.035`. I would recommend trying 0 first, then increasing by `0.01` increments to find the sweet spot, then use that value in TTA. \n\nAlso, make sure the testing data is scaled the same as the training data when applying noise TTA. You want to replicate training augmentation as much as possible.",
    "3063988": "Maybe… Let’s what happen going on!"
  },
  "source": "meta"
}