{
  "id": 395868,
  "title": "[place holder] stable diffusion for extra data generation ",
  "url": "/competitions/asl-signs/discussion/395868",
  "author_name": "",
  "post_date": "2023-03-19T08:11:49.492859400Z",
  "votes": 16,
  "comment_count": 5,
  "views": 0,
  "content": "<p><strong>DISCLAIMER !!!: The following is an exploratory idea and it may not work. If your goal is to achieve a high medal ranking, it is recommended to focus on the conventional model engineering and hyper-parameter tuning methods before trying this approach.</strong></p>\n<p>OBJECTIVE: <br>\nIn 2023, AI, particularly generative AI, is expected to explode. Its core technologies are generative pretraining (GPT), reinforcement learning from human feedback (RLHF), and probabilistic denoising diffusion model (DDPM). Hence, it may be worthwhile to explore these new technologies.</p>\n<p>The objective of this exploration is to investigate the use of DDPM to generate new training data, whether labeled or unlabeled. This is a complex topic. I will explain the basic (but not the very basic), so it's recommended for Kagglers to do some background reading beforehand. Here are some resources:</p>\n<ul>\n<li>A 1D example: <br>\n<a href=\"https://www.kaggle.com/code/grishasizov/simple-denoising-diffusion-model-toy-1d-example\" target=\"_blank\">https://www.kaggle.com/code/grishasizov/simple-denoising-diffusion-model-toy-1d-example</a>  </li>\n<li>A 2D example (set of 2D points): <br>\n<a href=\"https://github.com/tanelp/tiny-diffusion\" target=\"_blank\">https://github.com/tanelp/tiny-diffusion</a>    </li>\n<li>An introduction to the classic paper \"Denoising Diffusion Probabilistic Models\":   <br>\n<a href=\"https://www.youtube.com/watch?v=a4Yfz2FxXiY\" target=\"_blank\">https://www.youtube.com/watch?v=a4Yfz2FxXiY</a> and <a href=\"https://arxiv.org/pdf/2006.11239.pdf\" target=\"_blank\">https://arxiv.org/pdf/2006.11239.pdf</a>  </li>\n</ul>\n<p>It's important to note that this is a deep topic, and the aim of this exploration is to test the feasibility of using DDPM for generating training data. Therefore, it's not a guarantee of success or high medal rankings.</p>\n<p>some work using DDPM to generate skeletion landmaks :<br>\n… to be updated …</p>\n<p>[1] MDM: Human Motion Diffusion Model<br>\n<a href=\"https://guytevet.github.io/mdm-page/\" target=\"_blank\">https://guytevet.github.io/mdm-page/</a></p>\n<p>[2] Taming Diffusion Models for Audio-Driven Co-Speech Gesture Generation<br>\n<a href=\"https://arxiv.org/pdf/2303.09119.pdf\" target=\"_blank\">https://arxiv.org/pdf/2303.09119.pdf</a></p>\n<p><img src=\"https://i.ibb.co/1Rvdd6W/Selection-999-1537.png\" alt=\"https://i.ibb.co/1Rvdd6W/Selection-999-1537.png\"></p>\n<hr>\n<p>work in progress … will take about a few weeks to complete with notebook and experiment results … </p>\n<p>(post improved by chatgpt3)</p>",
  "messages": [
    {
      "id": "2187987",
      "postDate": "03/19/2023 08:11:49",
      "content": "<p><strong>DISCLAIMER !!!: The following is an exploratory idea and it may not work. If your goal is to achieve a high medal ranking, it is recommended to focus on the conventional model engineering and hyper-parameter tuning methods before trying this approach.</strong></p>\n<p>OBJECTIVE: <br>\nIn 2023, AI, particularly generative AI, is expected to explode. Its core technologies are generative pretraining (GPT), reinforcement learning from human feedback (RLHF), and probabilistic denoising diffusion model (DDPM). Hence, it may be worthwhile to explore these new technologies.</p>\n<p>The objective of this exploration is to investigate the use of DDPM to generate new training data, whether labeled or unlabeled. This is a complex topic. I will explain the basic (but not the very basic), so it's recommended for Kagglers to do some background reading beforehand. Here are some resources:</p>\n<ul>\n<li>A 1D example: <br>\n<a href=\"https://www.kaggle.com/code/grishasizov/simple-denoising-diffusion-model-toy-1d-example\" target=\"_blank\">https://www.kaggle.com/code/grishasizov/simple-denoising-diffusion-model-toy-1d-example</a>  </li>\n<li>A 2D example (set of 2D points): <br>\n<a href=\"https://github.com/tanelp/tiny-diffusion\" target=\"_blank\">https://github.com/tanelp/tiny-diffusion</a>    </li>\n<li>An introduction to the classic paper \"Denoising Diffusion Probabilistic Models\":   <br>\n<a href=\"https://www.youtube.com/watch?v=a4Yfz2FxXiY\" target=\"_blank\">https://www.youtube.com/watch?v=a4Yfz2FxXiY</a> and <a href=\"https://arxiv.org/pdf/2006.11239.pdf\" target=\"_blank\">https://arxiv.org/pdf/2006.11239.pdf</a>  </li>\n</ul>\n<p>It's important to note that this is a deep topic, and the aim of this exploration is to test the feasibility of using DDPM for generating training data. Therefore, it's not a guarantee of success or high medal rankings.</p>\n<p>some work using DDPM to generate skeletion landmaks :<br>\n… to be updated …</p>\n<p>[1] MDM: Human Motion Diffusion Model<br>\n<a href=\"https://guytevet.github.io/mdm-page/\" target=\"_blank\">https://guytevet.github.io/mdm-page/</a></p>\n<p>[2] Taming Diffusion Models for Audio-Driven Co-Speech Gesture Generation<br>\n<a href=\"https://arxiv.org/pdf/2303.09119.pdf\" target=\"_blank\">https://arxiv.org/pdf/2303.09119.pdf</a></p>\n<p><img src=\"https://i.ibb.co/1Rvdd6W/Selection-999-1537.png\" alt=\"https://i.ibb.co/1Rvdd6W/Selection-999-1537.png\"></p>\n<hr>\n<p>work in progress … will take about a few weeks to complete with notebook and experiment results … </p>\n<p>(post improved by chatgpt3)</p>",
      "rawMarkdown": "**DISCLAIMER !!!: The following is an exploratory idea and it may not work. If your goal is to achieve a high medal ranking, it is recommended to focus on the conventional model engineering and hyper-parameter tuning methods before trying this approach.**\n\nOBJECTIVE: \nIn 2023, AI, particularly generative AI, is expected to explode. Its core technologies are generative pretraining (GPT), reinforcement learning from human feedback (RLHF), and probabilistic denoising diffusion model (DDPM). Hence, it may be worthwhile to explore these new technologies.\n\nThe objective of this exploration is to investigate the use of DDPM to generate new training data, whether labeled or unlabeled. This is a complex topic. I will explain the basic (but not the very basic), so it's recommended for Kagglers to do some background reading beforehand. Here are some resources:\n\n- A 1D example: \nhttps://www.kaggle.com/code/grishasizov/simple-denoising-diffusion-model-toy-1d-example  \n- A 2D example (set of 2D points): \nhttps://github.com/tanelp/tiny-diffusion    \n- An introduction to the classic paper \"Denoising Diffusion Probabilistic Models\":   \nhttps://www.youtube.com/watch?v=a4Yfz2FxXiY and https://arxiv.org/pdf/2006.11239.pdf  \n\n\nIt's important to note that this is a deep topic, and the aim of this exploration is to test the feasibility of using DDPM for generating training data. Therefore, it's not a guarantee of success or high medal rankings.\n\n\nsome work using DDPM to generate skeletion landmaks :\n... to be updated ...\n\n[1] MDM: Human Motion Diffusion Model\nhttps://guytevet.github.io/mdm-page/\n\n\n[2] Taming Diffusion Models for Audio-Driven Co-Speech Gesture Generation\nhttps://arxiv.org/pdf/2303.09119.pdf\n\n![https://i.ibb.co/1Rvdd6W/Selection-999-1537.png](https://i.ibb.co/1Rvdd6W/Selection-999-1537.png)\n\n\n----\n\nwork in progress ... will take about a few weeks to complete with notebook and experiment results ... \n\n\n(post improved by chatgpt3)",
      "votes": null
    },
    {
      "id": "2192066",
      "postDate": "03/22/2023 12:03:53",
      "content": "<p>Maybe the generated data is not good enough for the cls model to train on but we can try MAE on that.</p>",
      "rawMarkdown": "Maybe the generated data is not good enough for the cls model to train on but we can try MAE on that.",
      "votes": null
    },
    {
      "id": "2195611",
      "postDate": "03/24/2023 18:48:54",
      "content": "<p>amazing AI application</p>",
      "rawMarkdown": "amazing AI application",
      "votes": null
    },
    {
      "id": "2201036",
      "postDate": "03/29/2023 02:12:16",
      "content": "<p>if we can do something like this:<br>\nGestureDiffuCLIP: Gesture Diffusion Model with CLIP Latents<br>\n<a href=\"https://arxiv.org/abs/2303.14613\" target=\"_blank\">https://arxiv.org/abs/2303.14613</a></p>\n<p>movie:<br>\n<a href=\"https://www.youtube.com/watch?v=Psi1IOZGq8c\" target=\"_blank\">https://www.youtube.com/watch?v=Psi1IOZGq8c</a></p>",
      "rawMarkdown": "if we can do something like this:\nGestureDiffuCLIP: Gesture Diffusion Model with CLIP Latents\nhttps://arxiv.org/abs/2303.14613\n\nmovie:\nhttps://www.youtube.com/watch?v=Psi1IOZGq8c",
      "votes": null
    },
    {
      "id": "2202380",
      "postDate": "03/30/2023 02:25:30",
      "content": "<p>instead of generating more \"signer\", how about learning a diffusion model to convert a given xyz to a canoical signer?<br>\n(aka normalisation)</p>",
      "rawMarkdown": "instead of generating more \"signer\", how about learning a diffusion model to convert a given xyz to a canoical signer?\n(aka normalisation)",
      "votes": null
    },
    {
      "id": "2206483",
      "postDate": "04/02/2023 15:31:29",
      "content": "<p>GAN generated video: <a href=\"https://www.youtube.com/watch?v=wOxWUyXX6Ys\" target=\"_blank\">https://www.youtube.com/watch?v=wOxWUyXX6Ys</a></p>",
      "rawMarkdown": "GAN generated video: https://www.youtube.com/watch?v=wOxWUyXX6Ys",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2192066,
      "author_name": "chinartist",
      "author_url": "",
      "post_date": "03/22/2023 12:03:53",
      "content": "<p>Maybe the generated data is not good enough for the cls model to train on but we can try MAE on that.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2195611,
      "author_name": "dzer77",
      "author_url": "",
      "post_date": "03/24/2023 18:48:54",
      "content": "<p>amazing AI application</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2201036,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/29/2023 02:12:16",
      "content": "<p>if we can do something like this:<br>\nGestureDiffuCLIP: Gesture Diffusion Model with CLIP Latents<br>\n<a href=\"https://arxiv.org/abs/2303.14613\" target=\"_blank\">https://arxiv.org/abs/2303.14613</a></p>\n<p>movie:<br>\n<a href=\"https://www.youtube.com/watch?v=Psi1IOZGq8c\" target=\"_blank\">https://www.youtube.com/watch?v=Psi1IOZGq8c</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2202380,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/30/2023 02:25:30",
      "content": "<p>instead of generating more \"signer\", how about learning a diffusion model to convert a given xyz to a canoical signer?<br>\n(aka normalisation)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2206483,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/02/2023 15:31:29",
      "content": "<p>GAN generated video: <a href=\"https://www.youtube.com/watch?v=wOxWUyXX6Ys\" target=\"_blank\">https://www.youtube.com/watch?v=wOxWUyXX6Ys</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2187987": "**DISCLAIMER !!!: The following is an exploratory idea and it may not work. If your goal is to achieve a high medal ranking, it is recommended to focus on the conventional model engineering and hyper-parameter tuning methods before trying this approach.**\n\nOBJECTIVE: \nIn 2023, AI, particularly generative AI, is expected to explode. Its core technologies are generative pretraining (GPT), reinforcement learning from human feedback (RLHF), and probabilistic denoising diffusion model (DDPM). Hence, it may be worthwhile to explore these new technologies.\n\nThe objective of this exploration is to investigate the use of DDPM to generate new training data, whether labeled or unlabeled. This is a complex topic. I will explain the basic (but not the very basic), so it's recommended for Kagglers to do some background reading beforehand. Here are some resources:\n\n- A 1D example: \nhttps://www.kaggle.com/code/grishasizov/simple-denoising-diffusion-model-toy-1d-example  \n- A 2D example (set of 2D points): \nhttps://github.com/tanelp/tiny-diffusion    \n- An introduction to the classic paper \"Denoising Diffusion Probabilistic Models\":   \nhttps://www.youtube.com/watch?v=a4Yfz2FxXiY and https://arxiv.org/pdf/2006.11239.pdf  \n\n\nIt's important to note that this is a deep topic, and the aim of this exploration is to test the feasibility of using DDPM for generating training data. Therefore, it's not a guarantee of success or high medal rankings.\n\n\nsome work using DDPM to generate skeletion landmaks :\n... to be updated ...\n\n[1] MDM: Human Motion Diffusion Model\nhttps://guytevet.github.io/mdm-page/\n\n\n[2] Taming Diffusion Models for Audio-Driven Co-Speech Gesture Generation\nhttps://arxiv.org/pdf/2303.09119.pdf\n\n![https://i.ibb.co/1Rvdd6W/Selection-999-1537.png](https://i.ibb.co/1Rvdd6W/Selection-999-1537.png)\n\n\n----\n\nwork in progress ... will take about a few weeks to complete with notebook and experiment results ... \n\n\n(post improved by chatgpt3)",
    "2192066": "Maybe the generated data is not good enough for the cls model to train on but we can try MAE on that.",
    "2195611": "amazing AI application",
    "2201036": "if we can do something like this:\nGestureDiffuCLIP: Gesture Diffusion Model with CLIP Latents\nhttps://arxiv.org/abs/2303.14613\n\nmovie:\nhttps://www.youtube.com/watch?v=Psi1IOZGq8c",
    "2202380": "instead of generating more \"signer\", how about learning a diffusion model to convert a given xyz to a canoical signer?\n(aka normalisation)",
    "2206483": "GAN generated video: https://www.youtube.com/watch?v=wOxWUyXX6Ys"
  },
  "source": "meta"
}