{
  "id": 44239,
  "title": "Is all synthetic data allowed?",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/44239",
  "author_name": "",
  "post_date": "2017-11-25T19:35:31.522215700Z",
  "votes": 21,
  "comment_count": 9,
  "views": 0,
  "content": "<p>1) Can we augment data/create synthetic data of any kind as long as the we use nothing but the competition data to base it on?</p>\n\n<p>2) What about generating synthetic noise data, which is just based off of the definition of the noise?</p>\n\n<p>3) While its been asked elsewhere: Does it matter whether we use only the train set for creating synthetic data, or can we use both the train and test set, as long as no human labeling, or labeling technology that uses outside information is used?</p>",
  "messages": [
    {
      "id": "248351",
      "postDate": "11/25/2017 19:35:31",
      "content": "<p>1) Can we augment data/create synthetic data of any kind as long as the we use nothing but the competition data to base it on?</p>\n\n<p>2) What about generating synthetic noise data, which is just based off of the definition of the noise?</p>\n\n<p>3) While its been asked elsewhere: Does it matter whether we use only the train set for creating synthetic data, or can we use both the train and test set, as long as no human labeling, or labeling technology that uses outside information is used?</p>",
      "rawMarkdown": "1) Can we augment data/create synthetic data of any kind as long as the we use nothing but the competition data to base it on?\n\n2) What about generating synthetic noise data, which is just based off of the definition of the noise?\n\n3) While its been asked elsewhere: Does it matter whether we use only the train set for creating synthetic data, or can we use both the train and test set, as long as no human labeling, or labeling technology that uses outside information is used?",
      "votes": null
    },
    {
      "id": "249418",
      "postDate": "11/28/2017 12:30:24",
      "content": "<p>I'm also curious about this part.</p>",
      "rawMarkdown": "I'm also curious about this part.",
      "votes": null
    },
    {
      "id": "250153",
      "postDate": "11/29/2017 21:49:32",
      "content": "<p>The pink and white noise provided is synthetic, and they give us the code how it was generated. I doubt there's an issue with generating your own.</p>",
      "rawMarkdown": "The pink and white noise provided is synthetic, and they give us the code how it was generated. I doubt there's an issue with generating your own.",
      "votes": null
    },
    {
      "id": "253938",
      "postDate": "12/05/2017 22:01:26",
      "content": "<blockquote>\n  <p>1) Can we augment data/create synthetic data of any kind as long as\n  we use nothing but the competition data to base it on?</p>\n</blockquote>\n\n<p>Using techniques to augment the training data (for example adding noise, pitch shifting) should be fine, as long as the changes are purely algorithmic and not using any captured data (or anything based on captured data, like pretrained models).</p>\n\n<blockquote>\n  <p>2) What about generating synthetic noise data, which is just based off\n  of the definition of the noise?</p>\n</blockquote>\n\n<p>Synthetic noise data that's algorithmically generated is fine to use.</p>\n\n<blockquote>\n  <p>3) While its been asked elsewhere: Does it matter whether we use only\n  the train set for creating synthetic data, or can we use both the\n  train and test set, as long as no human labeling, or labeling\n  technology that uses outside information is used?</p>\n</blockquote>\n\n<p>I am not the final word on this (I would ping @inversion), but I believe only the training set should be used as inputs.</p>",
      "rawMarkdown": "&gt; 1) Can we augment data/create synthetic data of any kind as long as\n&gt; we use nothing but the competition data to base it on?\n\nUsing techniques to augment the training data (for example adding noise, pitch shifting) should be fine, as long as the changes are purely algorithmic and not using any captured data (or anything based on captured data, like pretrained models).\n\n&gt; 2) What about generating synthetic noise data, which is just based off\n&gt; of the definition of the noise?\n\nSynthetic noise data that's algorithmically generated is fine to use.\n\n&gt; 3) While its been asked elsewhere: Does it matter whether we use only\n&gt; the train set for creating synthetic data, or can we use both the\n&gt; train and test set, as long as no human labeling, or labeling\n&gt; technology that uses outside information is used?\n\nI am not the final word on this (I would ping @inversion), but I believe only the training set should be used as inputs.",
      "votes": null
    },
    {
      "id": "254309",
      "postDate": "12/06/2017 17:54:06",
      "content": "<blockquote>\n  <p>I am not the final word on this (I would ping @inversion), but I believe only the training set should be used as inputs.</p>\n</blockquote>\n\n<p>Knowing the intention is good enough. If there is salt in the test set it might be difficult to use some unsupervised methods anyway. </p>\n\n<p>Thanks for the clarification! </p>",
      "rawMarkdown": "&gt; I am not the final word on this (I would ping @inversion), but I believe only the training set should be used as inputs.\n\nKnowing the intention is good enough. If there is salt in the test set it might be difficult to use some unsupervised methods anyway. \n\nThanks for the clarification!",
      "votes": null
    },
    {
      "id": "254368",
      "postDate": "12/06/2017 20:14:27",
      "content": "<p>Hi phalaris - Yes, in general, unsupervised learning is fine with the Test set. As long as, per the Rules, there is no hand labeling of the Test set.</p>",
      "rawMarkdown": "Hi phalaris - Yes, in general, unsupervised learning is fine with the Test set. As long as, per the Rules, there is no hand labeling of the Test set.",
      "votes": null
    },
    {
      "id": "254706",
      "postDate": "12/07/2017 13:01:35",
      "content": "<p>So, just to be 100% sure. A combination of supervised learning on the training data set + unsupervised learning on the test data set is allowed, correct?</p>",
      "rawMarkdown": "So, just to be 100% sure. A combination of supervised learning on the training data set + unsupervised learning on the test data set is allowed, correct?",
      "votes": null
    },
    {
      "id": "254741",
      "postDate": "12/07/2017 14:34:46",
      "content": "<p>Correct.</p>",
      "rawMarkdown": "Correct.",
      "votes": null
    },
    {
      "id": "254924",
      "postDate": "12/07/2017 21:59:29",
      "content": "<p>That changes the game.</p>\n\n<p>Like @Pete, my feeling is that only training data should be used.</p>\n\n<p>By using test data, we can end up with models that scores high on Kaggle but not on real world. An unlabeled training dataset should have been provided to avoid that.</p>\n\n<p>Now, a suggestion for everyone for semi supervised learning is this recently published paper on NIPS 2017:</p>\n\n<p><a href=\"https://arxiv.org/abs/1711.07476v1\">https://arxiv.org/abs/1711.07476v1</a></p>",
      "rawMarkdown": "That changes the game.\n\nLike @Pete, my feeling is that only training data should be used.\n\nBy using test data, we can end up with models that scores high on Kaggle but not on real world. An unlabeled training dataset should have been provided to avoid that.\n\nNow, a suggestion for everyone for semi supervised learning is this recently published paper on NIPS 2017:\n\nhttps://arxiv.org/abs/1711.07476v1",
      "votes": null
    },
    {
      "id": "256193",
      "postDate": "12/11/2017 12:46:20",
      "content": "<p>You've got a good point. We might end up overfitting the data by incorporating the unlabeled test data set as part of our training data set and nothing else to validate against.</p>\n\n<p>I've read about ladder for RNNs in the past, but the idea of virtual adversarial noise is very new to me. Thanks for the link. I'll give it a shot later.</p>\n\n<p>Boa sorte!</p>",
      "rawMarkdown": "You've got a good point. We might end up overfitting the data by incorporating the unlabeled test data set as part of our training data set and nothing else to validate against.\n\nI've read about ladder for RNNs in the past, but the idea of virtual adversarial noise is very new to me. Thanks for the link. I'll give it a shot later.\n\nBoa sorte!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 249418,
      "author_name": "jongpillee",
      "author_url": "",
      "post_date": "11/28/2017 12:30:24",
      "content": "<p>I'm also curious about this part.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 250153,
      "author_name": "gerti004",
      "author_url": "",
      "post_date": "11/29/2017 21:49:32",
      "content": "<p>The pink and white noise provided is synthetic, and they give us the code how it was generated. I doubt there's an issue with generating your own.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 253938,
      "author_name": "petewarden",
      "author_url": "",
      "post_date": "12/05/2017 22:01:26",
      "content": "<blockquote>\n  <p>1) Can we augment data/create synthetic data of any kind as long as\n  we use nothing but the competition data to base it on?</p>\n</blockquote>\n\n<p>Using techniques to augment the training data (for example adding noise, pitch shifting) should be fine, as long as the changes are purely algorithmic and not using any captured data (or anything based on captured data, like pretrained models).</p>\n\n<blockquote>\n  <p>2) What about generating synthetic noise data, which is just based off\n  of the definition of the noise?</p>\n</blockquote>\n\n<p>Synthetic noise data that's algorithmically generated is fine to use.</p>\n\n<blockquote>\n  <p>3) While its been asked elsewhere: Does it matter whether we use only\n  the train set for creating synthetic data, or can we use both the\n  train and test set, as long as no human labeling, or labeling\n  technology that uses outside information is used?</p>\n</blockquote>\n\n<p>I am not the final word on this (I would ping @inversion), but I believe only the training set should be used as inputs.</p>",
      "votes": null,
      "replies": [
        {
          "id": 254309,
          "author_name": "devinanzelmo",
          "author_url": "",
          "post_date": "12/06/2017 17:54:06",
          "content": "<blockquote>\n  <p>I am not the final word on this (I would ping @inversion), but I believe only the training set should be used as inputs.</p>\n</blockquote>\n\n<p>Knowing the intention is good enough. If there is salt in the test set it might be difficult to use some unsupervised methods anyway. </p>\n\n<p>Thanks for the clarification! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 254368,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "12/06/2017 20:14:27",
          "content": "<p>Hi phalaris - Yes, in general, unsupervised learning is fine with the Test set. As long as, per the Rules, there is no hand labeling of the Test set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 254706,
          "author_name": "ekonishi",
          "author_url": "",
          "post_date": "12/07/2017 13:01:35",
          "content": "<p>So, just to be 100% sure. A combination of supervised learning on the training data set + unsupervised learning on the test data set is allowed, correct?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 254741,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "12/07/2017 14:34:46",
          "content": "<p>Correct.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 254924,
          "author_name": "barbolo",
          "author_url": "",
          "post_date": "12/07/2017 21:59:29",
          "content": "<p>That changes the game.</p>\n\n<p>Like @Pete, my feeling is that only training data should be used.</p>\n\n<p>By using test data, we can end up with models that scores high on Kaggle but not on real world. An unlabeled training dataset should have been provided to avoid that.</p>\n\n<p>Now, a suggestion for everyone for semi supervised learning is this recently published paper on NIPS 2017:</p>\n\n<p><a href=\"https://arxiv.org/abs/1711.07476v1\">https://arxiv.org/abs/1711.07476v1</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 256193,
          "author_name": "ekonishi",
          "author_url": "",
          "post_date": "12/11/2017 12:46:20",
          "content": "<p>You've got a good point. We might end up overfitting the data by incorporating the unlabeled test data set as part of our training data set and nothing else to validate against.</p>\n\n<p>I've read about ladder for RNNs in the past, but the idea of virtual adversarial noise is very new to me. Thanks for the link. I'll give it a shot later.</p>\n\n<p>Boa sorte!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "248351": "1) Can we augment data/create synthetic data of any kind as long as the we use nothing but the competition data to base it on?\n\n2) What about generating synthetic noise data, which is just based off of the definition of the noise?\n\n3) While its been asked elsewhere: Does it matter whether we use only the train set for creating synthetic data, or can we use both the train and test set, as long as no human labeling, or labeling technology that uses outside information is used?",
    "249418": "I'm also curious about this part.",
    "250153": "The pink and white noise provided is synthetic, and they give us the code how it was generated. I doubt there's an issue with generating your own.",
    "253938": "&gt; 1) Can we augment data/create synthetic data of any kind as long as\n&gt; we use nothing but the competition data to base it on?\n\nUsing techniques to augment the training data (for example adding noise, pitch shifting) should be fine, as long as the changes are purely algorithmic and not using any captured data (or anything based on captured data, like pretrained models).\n\n&gt; 2) What about generating synthetic noise data, which is just based off\n&gt; of the definition of the noise?\n\nSynthetic noise data that's algorithmically generated is fine to use.\n\n&gt; 3) While its been asked elsewhere: Does it matter whether we use only\n&gt; the train set for creating synthetic data, or can we use both the\n&gt; train and test set, as long as no human labeling, or labeling\n&gt; technology that uses outside information is used?\n\nI am not the final word on this (I would ping @inversion), but I believe only the training set should be used as inputs.",
    "254309": "&gt; I am not the final word on this (I would ping @inversion), but I believe only the training set should be used as inputs.\n\nKnowing the intention is good enough. If there is salt in the test set it might be difficult to use some unsupervised methods anyway. \n\nThanks for the clarification!",
    "254368": "Hi phalaris - Yes, in general, unsupervised learning is fine with the Test set. As long as, per the Rules, there is no hand labeling of the Test set.",
    "254706": "So, just to be 100% sure. A combination of supervised learning on the training data set + unsupervised learning on the test data set is allowed, correct?",
    "254741": "Correct.",
    "254924": "That changes the game.\n\nLike @Pete, my feeling is that only training data should be used.\n\nBy using test data, we can end up with models that scores high on Kaggle but not on real world. An unlabeled training dataset should have been provided to avoid that.\n\nNow, a suggestion for everyone for semi supervised learning is this recently published paper on NIPS 2017:\n\nhttps://arxiv.org/abs/1711.07476v1",
    "256193": "You've got a good point. We might end up overfitting the data by incorporating the unlabeled test data set as part of our training data set and nothing else to validate against.\n\nI've read about ladder for RNNs in the past, but the idea of virtual adversarial noise is very new to me. Thanks for the link. I'll give it a shot later.\n\nBoa sorte!"
  },
  "source": "meta"
}