{
  "id": 557839,
  "title": "Synthetic data and original data",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/557839",
  "author_name": "",
  "post_date": "2025-01-21T14:59:15.062129100Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I have tried using synthetic data for verification and found that the results are better than the original data, why?<br>\nThe training results of the original data are not very good.</p>",
  "messages": [
    {
      "id": "3101998",
      "postDate": "01/21/2025 14:59:15",
      "content": "<p>I have tried using synthetic data for verification and found that the results are better than the original data, why?<br>\nThe training results of the original data are not very good.</p>",
      "rawMarkdown": "I have tried using synthetic data for verification and found that the results are better than the original data, why?\nThe training results of the original data are not very good.",
      "votes": null
    },
    {
      "id": "3109221",
      "postDate": "01/28/2025 16:14:12",
      "content": "<p>Can you give a bit more information about what you are referring to?  Are you saying that you get better training results when training on the synthetic data?  </p>\n<p>Keep in mind the synthetic data and competition training data are very different.  We are working with denoised real data in teh competition that better reflects the unseen test set whilst the synthetic data does not follow the same distribution.  I recommend pretraining on the synthetic data and then fine-tune with the competition data.  This will give you a better starting point but will make sure your model is trained on data more similar to the unseen test data.</p>",
      "rawMarkdown": "Can you give a bit more information about what you are referring to?  Are you saying that you get better training results when training on the synthetic data?  \n\nKeep in mind the synthetic data and competition training data are very different.  We are working with denoised real data in teh competition that better reflects the unseen test set whilst the synthetic data does not follow the same distribution.  I recommend pretraining on the synthetic data and then fine-tune with the competition data.  This will give you a better starting point but will make sure your model is trained on data more similar to the unseen test data.",
      "votes": null
    },
    {
      "id": "3109222",
      "postDate": "01/28/2025 16:14:30",
      "content": "<p>I could be misunderstanding what you are talking about so please let me know!</p>",
      "rawMarkdown": "I could be misunderstanding what you are talking about so please let me know!",
      "votes": null
    },
    {
      "id": "3109509",
      "postDate": "01/29/2025 00:31:42",
      "content": "<p>You can see the details of these two notebooks should be able to understand my meaning</p>\n<p><a href=\"https://www.kaggle.com/code/zyh1104/czii-yolo-original-synthetic-data-ensemble\" target=\"_blank\">YOLO notebook</a></p>\n<p><a href=\"https://www.kaggle.com/code/zyh1104/czii-yolo11-original-synthetic-data-unet3d\" target=\"_blank\">YOLO+Unet3D notebook</a></p>",
      "rawMarkdown": "You can see the details of these two notebooks should be able to understand my meaning\n\n[YOLO notebook](https://www.kaggle.com/code/zyh1104/czii-yolo-original-synthetic-data-ensemble)\n\n[YOLO+Unet3D notebook](https://www.kaggle.com/code/zyh1104/czii-yolo11-original-synthetic-data-unet3d)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3109221,
      "author_name": "connorjd",
      "author_url": "",
      "post_date": "01/28/2025 16:14:12",
      "content": "<p>Can you give a bit more information about what you are referring to?  Are you saying that you get better training results when training on the synthetic data?  </p>\n<p>Keep in mind the synthetic data and competition training data are very different.  We are working with denoised real data in teh competition that better reflects the unseen test set whilst the synthetic data does not follow the same distribution.  I recommend pretraining on the synthetic data and then fine-tune with the competition data.  This will give you a better starting point but will make sure your model is trained on data more similar to the unseen test data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3109222,
          "author_name": "connorjd",
          "author_url": "",
          "post_date": "01/28/2025 16:14:30",
          "content": "<p>I could be misunderstanding what you are talking about so please let me know!</p>",
          "votes": null,
          "replies": [
            {
              "id": 3109509,
              "author_name": "zyh1104",
              "author_url": "",
              "post_date": "01/29/2025 00:31:42",
              "content": "<p>You can see the details of these two notebooks should be able to understand my meaning</p>\n<p><a href=\"https://www.kaggle.com/code/zyh1104/czii-yolo-original-synthetic-data-ensemble\" target=\"_blank\">YOLO notebook</a></p>\n<p><a href=\"https://www.kaggle.com/code/zyh1104/czii-yolo11-original-synthetic-data-unet3d\" target=\"_blank\">YOLO+Unet3D notebook</a></p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3101998": "I have tried using synthetic data for verification and found that the results are better than the original data, why?\nThe training results of the original data are not very good.",
    "3109221": "Can you give a bit more information about what you are referring to?  Are you saying that you get better training results when training on the synthetic data?  \n\nKeep in mind the synthetic data and competition training data are very different.  We are working with denoised real data in teh competition that better reflects the unseen test set whilst the synthetic data does not follow the same distribution.  I recommend pretraining on the synthetic data and then fine-tune with the competition data.  This will give you a better starting point but will make sure your model is trained on data more similar to the unseen test data.",
    "3109222": "I could be misunderstanding what you are talking about so please let me know!",
    "3109509": "You can see the details of these two notebooks should be able to understand my meaning\n\n[YOLO notebook](https://www.kaggle.com/code/zyh1104/czii-yolo-original-synthetic-data-ensemble)\n\n[YOLO+Unet3D notebook](https://www.kaggle.com/code/zyh1104/czii-yolo11-original-synthetic-data-unet3d)"
  },
  "source": "meta"
}