{
  "id": 541638,
  "title": "Generative Adversarial Networks for synthetic data",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/541638",
  "author_name": "",
  "post_date": "2024-10-20T16:21:27.239220Z",
  "votes": 3,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Hello everyone</p>\n<p>Dear fellows as we know that we have limited data so, I thought to increase the data by generating synthetic data by using GANs. You can view notebook <a href=\"https://www.kaggle.com/code/taimour/generative-adversarial-networks-lgb-piu-cmi\" target=\"_blank\">🕸️Generative Adversarial Networks | LGB | PIU CMI</a>.<br>\nAfter successfully generating data it gave QWK above 0.9 but when it came to LB or test data score was down to 0.171 <br>\nWhy it happened like this? Why a high QWK doesn't proves to be better for the test data?<br>\nDo I need to improve on the GAN further to improve the data? if yes then what changes are needed, any suggestions?</p>",
  "messages": [
    {
      "id": "3023508",
      "postDate": "10/20/2024 16:21:27",
      "content": "<p>Hello everyone</p>\n<p>Dear fellows as we know that we have limited data so, I thought to increase the data by generating synthetic data by using GANs. You can view notebook <a href=\"https://www.kaggle.com/code/taimour/generative-adversarial-networks-lgb-piu-cmi\" target=\"_blank\">🕸️Generative Adversarial Networks | LGB | PIU CMI</a>.<br>\nAfter successfully generating data it gave QWK above 0.9 but when it came to LB or test data score was down to 0.171 <br>\nWhy it happened like this? Why a high QWK doesn't proves to be better for the test data?<br>\nDo I need to improve on the GAN further to improve the data? if yes then what changes are needed, any suggestions?</p>",
      "rawMarkdown": "Hello everyone\n\nDear fellows as we know that we have limited data so, I thought to increase the data by generating synthetic data by using GANs. You can view notebook [🕸️Generative Adversarial Networks | LGB | PIU CMI](https://www.kaggle.com/code/taimour/generative-adversarial-networks-lgb-piu-cmi).\nAfter successfully generating data it gave QWK above 0.9 but when it came to LB or test data score was down to 0.171 \nWhy it happened like this? Why a high QWK doesn't proves to be better for the test data?\nDo I need to improve on the GAN further to improve the data? if yes then what changes are needed, any suggestions?",
      "votes": null
    },
    {
      "id": "3023705",
      "postDate": "10/20/2024 21:31:04",
      "content": "<p>I don't know GAN, but I guess your synthetic data is just twins of the real data. You have generated 8000 sii = 3 based on only 34 real cases (?) so you can imagine how that 8000 twins look like. And I guess the LGBM can easily learn to distinguish that 8000 synthetic sii 3 and boost QWK to 0.9.<br>\nOr we can say the test data is not similar to your synthetic data.<br>\nBut good try and keep your good work! Everyone here is trying to tell the model something extra without telling it anything.</p>",
      "rawMarkdown": "I don't know GAN, but I guess your synthetic data is just twins of the real data. You have generated 8000 sii = 3 based on only 34 real cases (?) so you can imagine how that 8000 twins look like. And I guess the LGBM can easily learn to distinguish that 8000 synthetic sii 3 and boost QWK to 0.9.\nOr we can say the test data is not similar to your synthetic data.\nBut good try and keep your good work! Everyone here is trying to tell the model something extra without telling it anything.",
      "votes": null
    },
    {
      "id": "3023810",
      "postDate": "10/21/2024 03:53:21",
      "content": "<p>Yes, I agree with you. It seems that test data is not similar to the generated synthetic data.</p>\n<p>As far as 8000 sii 3 is concerned I have tried with other combinations as well in which I had approx 2000+ sii for each 0,1,2 and less than 1000 sii 3. Even in that scenario I got QWK above 0.9</p>",
      "rawMarkdown": "Yes, I agree with you. It seems that test data is not similar to the generated synthetic data.\n\nAs far as 8000 sii 3 is concerned I have tried with other combinations as well in which I had approx 2000+ sii for each 0,1,2 and less than 1000 sii 3. Even in that scenario I got QWK above 0.9",
      "votes": null
    },
    {
      "id": "3024935",
      "postDate": "10/22/2024 07:01:58",
      "content": "<p>You might already be doing this, but make sure you perform augmentation within your training folds rather than before splitting the data. This way, you'll avoid having any augmented samples in your test fold, which could make your results too optimistic.</p>",
      "rawMarkdown": "You might already be doing this, but make sure you perform augmentation within your training folds rather than before splitting the data. This way, you'll avoid having any augmented samples in your test fold, which could make your results too optimistic.",
      "votes": null
    },
    {
      "id": "3024981",
      "postDate": "10/22/2024 08:08:42",
      "content": "<p>I find a question your volid data is from the train data  that has been transformed by GAN。SO you can try to split data before GAN model.</p>",
      "rawMarkdown": "I find a question your volid data is from the train data  that has been transformed by GAN。SO you can try to split data before GAN model.",
      "votes": null
    },
    {
      "id": "3025175",
      "postDate": "10/22/2024 13:51:46",
      "content": "<p>Ok, I will look into this. Thank you</p>",
      "rawMarkdown": "Ok, I will look into this. Thank you",
      "votes": null
    },
    {
      "id": "3025176",
      "postDate": "10/22/2024 13:52:23",
      "content": "<p>Sorry, I didn't understood. Can you please explain, thank you</p>",
      "rawMarkdown": "Sorry, I didn't understood. Can you please explain, thank you",
      "votes": null
    },
    {
      "id": "3031697",
      "postDate": "10/30/2024 02:29:22",
      "content": "<p>Seems like GAN has significantly distorted the distribution of the training set from the test, this is why it might seem like the model has severely over-fit. Have you tried applying the Box-Cox or Yeo-Johnson transformation to both the training and test sets to align their distributions, post GAIN?</p>",
      "rawMarkdown": "Seems like GAN has significantly distorted the distribution of the training set from the test, this is why it might seem like the model has severely over-fit. Have you tried applying the Box-Cox or Yeo-Johnson transformation to both the training and test sets to align their distributions, post GAIN?",
      "votes": null
    },
    {
      "id": "3031780",
      "postDate": "10/30/2024 05:19:39",
      "content": "<p>I have not applied power transformers up till now. I do plan to work further on GAN's in future and apply power transformers on it also. Thank you!</p>",
      "rawMarkdown": "I have not applied power transformers up till now. I do plan to work further on GAN's in future and apply power transformers on it also. Thank you!",
      "votes": null
    },
    {
      "id": "3036527",
      "postDate": "11/04/2024 16:31:34",
      "content": "<p><a href=\"https://www.kaggle.com/taimour\" target=\"_blank\">@taimour</a>  in genenaral  , usually even in tabular data , images, texts , that augment data using GANs, i am not sure but i think as we know GANs does not generate high good data , so the model; may detect some pattern that are not related to the real data but related to generator </p>",
      "rawMarkdown": "taimour  in genenaral  , usually even in tabular data , images, texts , that augment data using GANs, i am not sure but i think as we know GANs does not generate high good data , so the model; may detect some pattern that are not related to the real data but related to generator",
      "votes": null
    },
    {
      "id": "3036529",
      "postDate": "11/04/2024 16:34:31",
      "content": "<p><a href=\"https://www.kaggle.com/nigelxiang\" target=\"_blank\">@nigelxiang</a>  we are not use GANs to tranformed data , but to create and generate new data and concatenate with real data (train data ) in order to inscrease size of the data for taraining process</p>",
      "rawMarkdown": "nigelxiang  we are not use GANs to tranformed data , but to create and generate new data and concatenate with real data (train data ) in order to inscrease size of the data for taraining process",
      "votes": null
    },
    {
      "id": "3038233",
      "postDate": "11/06/2024 18:22:51",
      "content": "<p>Yes, I was just experimenting with different techniques to find out what results we can get. Based on the provided data we have to try different techniques to find any optimal solution. </p>",
      "rawMarkdown": "Yes, I was just experimenting with different techniques to find out what results we can get. Based on the provided data we have to try different techniques to find any optimal solution.",
      "votes": null
    },
    {
      "id": "3038236",
      "postDate": "11/06/2024 18:24:32",
      "content": "<p>You mean to say that I should generate data with GAN's before applying any kind of prepossessing on the data? Alright, thank you</p>",
      "rawMarkdown": "You mean to say that I should generate data with GAN's before applying any kind of prepossessing on the data? Alright, thank you",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3023705,
      "author_name": "tomyuen",
      "author_url": "",
      "post_date": "10/20/2024 21:31:04",
      "content": "<p>I don't know GAN, but I guess your synthetic data is just twins of the real data. You have generated 8000 sii = 3 based on only 34 real cases (?) so you can imagine how that 8000 twins look like. And I guess the LGBM can easily learn to distinguish that 8000 synthetic sii 3 and boost QWK to 0.9.<br>\nOr we can say the test data is not similar to your synthetic data.<br>\nBut good try and keep your good work! Everyone here is trying to tell the model something extra without telling it anything.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3023810,
          "author_name": "taimour",
          "author_url": "",
          "post_date": "10/21/2024 03:53:21",
          "content": "<p>Yes, I agree with you. It seems that test data is not similar to the generated synthetic data.</p>\n<p>As far as 8000 sii 3 is concerned I have tried with other combinations as well in which I had approx 2000+ sii for each 0,1,2 and less than 1000 sii 3. Even in that scenario I got QWK above 0.9</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3024935,
          "author_name": "welcomeworld",
          "author_url": "",
          "post_date": "10/22/2024 07:01:58",
          "content": "<p>You might already be doing this, but make sure you perform augmentation within your training folds rather than before splitting the data. This way, you'll avoid having any augmented samples in your test fold, which could make your results too optimistic.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3025175,
              "author_name": "taimour",
              "author_url": "",
              "post_date": "10/22/2024 13:51:46",
              "content": "<p>Ok, I will look into this. Thank you</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3024981,
      "author_name": "nigelxiang",
      "author_url": "",
      "post_date": "10/22/2024 08:08:42",
      "content": "<p>I find a question your volid data is from the train data  that has been transformed by GAN。SO you can try to split data before GAN model.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3025176,
          "author_name": "taimour",
          "author_url": "",
          "post_date": "10/22/2024 13:52:23",
          "content": "<p>Sorry, I didn't understood. Can you please explain, thank you</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3036529,
          "author_name": "saidkoussi",
          "author_url": "",
          "post_date": "11/04/2024 16:34:31",
          "content": "<p><a href=\"https://www.kaggle.com/nigelxiang\" target=\"_blank\">@nigelxiang</a>  we are not use GANs to tranformed data , but to create and generate new data and concatenate with real data (train data ) in order to inscrease size of the data for taraining process</p>",
          "votes": null,
          "replies": [
            {
              "id": 3038236,
              "author_name": "taimour",
              "author_url": "",
              "post_date": "11/06/2024 18:24:32",
              "content": "<p>You mean to say that I should generate data with GAN's before applying any kind of prepossessing on the data? Alright, thank you</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3031697,
      "author_name": "deepaksaldanha",
      "author_url": "",
      "post_date": "10/30/2024 02:29:22",
      "content": "<p>Seems like GAN has significantly distorted the distribution of the training set from the test, this is why it might seem like the model has severely over-fit. Have you tried applying the Box-Cox or Yeo-Johnson transformation to both the training and test sets to align their distributions, post GAIN?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3031780,
          "author_name": "taimour",
          "author_url": "",
          "post_date": "10/30/2024 05:19:39",
          "content": "<p>I have not applied power transformers up till now. I do plan to work further on GAN's in future and apply power transformers on it also. Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3036527,
      "author_name": "saidkoussi",
      "author_url": "",
      "post_date": "11/04/2024 16:31:34",
      "content": "<p><a href=\"https://www.kaggle.com/taimour\" target=\"_blank\">@taimour</a>  in genenaral  , usually even in tabular data , images, texts , that augment data using GANs, i am not sure but i think as we know GANs does not generate high good data , so the model; may detect some pattern that are not related to the real data but related to generator </p>",
      "votes": null,
      "replies": [
        {
          "id": 3038233,
          "author_name": "taimour",
          "author_url": "",
          "post_date": "11/06/2024 18:22:51",
          "content": "<p>Yes, I was just experimenting with different techniques to find out what results we can get. Based on the provided data we have to try different techniques to find any optimal solution. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3023508": "Hello everyone\n\nDear fellows as we know that we have limited data so, I thought to increase the data by generating synthetic data by using GANs. You can view notebook [🕸️Generative Adversarial Networks | LGB | PIU CMI](https://www.kaggle.com/code/taimour/generative-adversarial-networks-lgb-piu-cmi).\nAfter successfully generating data it gave QWK above 0.9 but when it came to LB or test data score was down to 0.171 \nWhy it happened like this? Why a high QWK doesn't proves to be better for the test data?\nDo I need to improve on the GAN further to improve the data? if yes then what changes are needed, any suggestions?",
    "3023705": "I don't know GAN, but I guess your synthetic data is just twins of the real data. You have generated 8000 sii = 3 based on only 34 real cases (?) so you can imagine how that 8000 twins look like. And I guess the LGBM can easily learn to distinguish that 8000 synthetic sii 3 and boost QWK to 0.9.\nOr we can say the test data is not similar to your synthetic data.\nBut good try and keep your good work! Everyone here is trying to tell the model something extra without telling it anything.",
    "3023810": "Yes, I agree with you. It seems that test data is not similar to the generated synthetic data.\n\nAs far as 8000 sii 3 is concerned I have tried with other combinations as well in which I had approx 2000+ sii for each 0,1,2 and less than 1000 sii 3. Even in that scenario I got QWK above 0.9",
    "3024935": "You might already be doing this, but make sure you perform augmentation within your training folds rather than before splitting the data. This way, you'll avoid having any augmented samples in your test fold, which could make your results too optimistic.",
    "3024981": "I find a question your volid data is from the train data  that has been transformed by GAN。SO you can try to split data before GAN model.",
    "3025175": "Ok, I will look into this. Thank you",
    "3025176": "Sorry, I didn't understood. Can you please explain, thank you",
    "3031697": "Seems like GAN has significantly distorted the distribution of the training set from the test, this is why it might seem like the model has severely over-fit. Have you tried applying the Box-Cox or Yeo-Johnson transformation to both the training and test sets to align their distributions, post GAIN?",
    "3031780": "I have not applied power transformers up till now. I do plan to work further on GAN's in future and apply power transformers on it also. Thank you!",
    "3036527": "taimour  in genenaral  , usually even in tabular data , images, texts , that augment data using GANs, i am not sure but i think as we know GANs does not generate high good data , so the model; may detect some pattern that are not related to the real data but related to generator",
    "3036529": "nigelxiang  we are not use GANs to tranformed data , but to create and generate new data and concatenate with real data (train data ) in order to inscrease size of the data for taraining process",
    "3038233": "Yes, I was just experimenting with different techniques to find out what results we can get. Based on the provided data we have to try different techniques to find any optimal solution.",
    "3038236": "You mean to say that I should generate data with GAN's before applying any kind of prepossessing on the data? Alright, thank you"
  },
  "source": "meta"
}