{
  "id": 544867,
  "title": "7 Training Samples / 500 Test Samples / 27 Extra Training Samples [Dataset] - How to trust CV with 7 samples? ",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/544867",
  "author_name": "",
  "post_date": "2024-11-07T10:38:14.652790300Z",
  "votes": 16,
  "comment_count": 6,
  "views": 0,
  "content": "<ul>\n<li>is it good to use training set as CV ?</li>\n<li>extra training samples as training ?</li>\n<li>500 test set !!!</li>\n</ul>\n<hr>\n<h1>How to download extra training?</h1>\n<pre><code> cryoet_data_portal  Client, Tomogram\n\nclient = Client()\n\ntomogram = Tomogram.get_by_id(client, ) \ntomogram.download_omezarr() \n</code></pre>\n<hr>\n<h1>Training # of objects - Average 181 objects</h1>\n<pre><code>Processing Run: TS_5_4\nAnnotating  objects ...\n\nProcessing Run: TS_69_2\nAnnotating  objects ...\n\nProcessing Run: TS_6_4\nAnnotating  objects ...\n\nProcessing Run: TS_6_6\nAnnotating  objects ...\n\nProcessing Run: TS_73_6\nAnnotating  objects ...\n\nProcessing Run: TS_86_3\nAnnotating  objects ...\n\nProcessing Run: TS_99_9\nAnnotating  objects ...\n</code></pre>\n<h1>From Github notebook - Average 1788 objects <a href=\"https://www.kaggle.com/datasets/seshurajup/czii-extra-dataset/data\" target=\"_blank\">Dataset - CZII Extra Dataset</a></h1>\n<pre><code>Processing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n</code></pre>",
  "messages": [
    {
      "id": "3038774",
      "postDate": "11/07/2024 10:38:14",
      "content": "<ul>\n<li>is it good to use training set as CV ?</li>\n<li>extra training samples as training ?</li>\n<li>500 test set !!!</li>\n</ul>\n<hr>\n<h1>How to download extra training?</h1>\n<pre><code> cryoet_data_portal  Client, Tomogram\n\nclient = Client()\n\ntomogram = Tomogram.get_by_id(client, ) \ntomogram.download_omezarr() \n</code></pre>\n<hr>\n<h1>Training # of objects - Average 181 objects</h1>\n<pre><code>Processing Run: TS_5_4\nAnnotating  objects ...\n\nProcessing Run: TS_69_2\nAnnotating  objects ...\n\nProcessing Run: TS_6_4\nAnnotating  objects ...\n\nProcessing Run: TS_6_6\nAnnotating  objects ...\n\nProcessing Run: TS_73_6\nAnnotating  objects ...\n\nProcessing Run: TS_86_3\nAnnotating  objects ...\n\nProcessing Run: TS_99_9\nAnnotating  objects ...\n</code></pre>\n<h1>From Github notebook - Average 1788 objects <a href=\"https://www.kaggle.com/datasets/seshurajup/czii-extra-dataset/data\" target=\"_blank\">Dataset - CZII Extra Dataset</a></h1>\n<pre><code>Processing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n\nProcessing Run: \nAnnotating  objects ...\n</code></pre>",
      "rawMarkdown": "* is it good to use training set as CV ?\n* extra training samples as training ?\n* 500 test set !!!\n\n----\n\n# How to download extra training?\n```python\nfrom cryoet_data_portal import Client, Tomogram\n\nclient = Client()\n\ntomogram = Tomogram.get_by_id(client, 17049) # - 16172 to 16198\ntomogram.download_omezarr() # many times its freeze and too slow. is good host provide these zarr also as external kaggle dataset\n```\n\n----\n\n# Training # of objects - Average 181 objects\n```python\nProcessing Run: TS_5_4\nAnnotating 140 objects ...\n\nProcessing Run: TS_69_2\nAnnotating 143 objects ...\n\nProcessing Run: TS_6_4\nAnnotating 193 objects ...\n\nProcessing Run: TS_6_6\nAnnotating 143 objects ...\n\nProcessing Run: TS_73_6\nAnnotating 217 objects ...\n\nProcessing Run: TS_86_3\nAnnotating 225 objects ...\n\nProcessing Run: TS_99_9\nAnnotating 208 objects ...\n```\n\n# From Github notebook - Average 1788 objects [Dataset - CZII Extra Dataset](https://www.kaggle.com/datasets/seshurajup/czii-extra-dataset/data)\n```python\nProcessing Run: 16172\nAnnotating 1774 objects ...\n\nProcessing Run: 16173\nAnnotating 1786 objects ...\n\nProcessing Run: 16174\nAnnotating 1784 objects ...\n\nProcessing Run: 16175\nAnnotating 1784 objects ...\n\nProcessing Run: 16176\nAnnotating 1781 objects ...\n\nProcessing Run: 16177\nAnnotating 1827 objects ...\n\nProcessing Run: 16178\nAnnotating 1781 objects ...\n\nProcessing Run: 16179\nAnnotating 1811 objects ...\n\nProcessing Run: 16180\nAnnotating 1808 objects ...\n\nProcessing Run: 16181\nAnnotating 1786 objects ...\n\nProcessing Run: 16182\nAnnotating 1777 objects ...\n\nProcessing Run: 16183\nAnnotating 1805 objects ...\n\nProcessing Run: 16184\nAnnotating 1791 objects ...\n\nProcessing Run: 16185\nAnnotating 1806 objects ...\n\nProcessing Run: 16186\nAnnotating 1759 objects ...\n\nProcessing Run: 16187\nAnnotating 1809 objects ...\n\nProcessing Run: 16188\nAnnotating 1770 objects ...\n\nProcessing Run: 16189\nAnnotating 1765 objects ...\n\nProcessing Run: 16190\nAnnotating 1796 objects ...\n\nProcessing Run: 16191\nAnnotating 1801 objects ...\n\nProcessing Run: 16192\nAnnotating 1766 objects ...\n\nProcessing Run: 16193\nAnnotating 1789 objects ...\n\nProcessing Run: 16194\nAnnotating 1799 objects ...\n\nProcessing Run: 16195\nAnnotating 1784 objects ...\n\nProcessing Run: 16196\nAnnotating 1783 objects ...\n\nProcessing Run: 16197\nAnnotating 1824 objects ...\n\nProcessing Run: 16198\nAnnotating 1751 objects ...\n```",
      "votes": null
    },
    {
      "id": "3038900",
      "postDate": "11/07/2024 13:59:51",
      "content": "<p>I'll make a more detailed post about the limited dataset today. This will include some explanation about <em>why</em> we have provided such little training data.</p>\n<p>There are extra datasets available for training, including synthetic data that matches the material in the test/training images. We have had luck combining the synthetic and experimental data. </p>\n<p>I will make a point of adding more documentation about getting the extra training data. Thank you for pointing all this out!!!</p>",
      "rawMarkdown": "I'll make a more detailed post about the limited dataset today. This will include some explanation about *why* we have provided such little training data.\n\nThere are extra datasets available for training, including synthetic data that matches the material in the test/training images. We have had luck combining the synthetic and experimental data. \n\nI will make a point of adding more documentation about getting the extra training data. Thank you for pointing all this out!!!",
      "votes": null
    },
    {
      "id": "3038909",
      "postDate": "11/07/2024 14:06:31",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/kharrington\" target=\"_blank\">@kharrington</a> for the quick response about data.</p>",
      "rawMarkdown": "Thanks @kharrington for the quick response about data.",
      "votes": null
    },
    {
      "id": "3039539",
      "postDate": "11/08/2024 06:51:51",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a> ,<br>\nThankyou for providing this valuable code. <br>\nBut I was having some difficulty in understanding some things here. In the heading, you said 27 extra training samples and then you are saying average 1788 objects. I am having confusion in that. It would be helpful if you could please clarify this ?</p>",
      "rawMarkdown": "Hi @seshurajup ,\nThankyou for providing this valuable code. \nBut I was having some difficulty in understanding some things here. In the heading, you said 27 extra training samples and then you are saying average 1788 objects. I am having confusion in that. It would be helpful if you could please clarify this ?",
      "votes": null
    },
    {
      "id": "3058943",
      "postDate": "11/30/2024 06:51:16",
      "content": "<p><a href=\"https://www.kaggle.com/kharrington\" target=\"_blank\">@kharrington</a> </p>\n<p>How can I get extra training data?</p>\n<p>It will be useful. </p>\n<p>Did you create the post or documentation?</p>",
      "rawMarkdown": "kharrington \n\nHow can I get extra training data?\n\nIt will be useful. \n\nDid you create the post or documentation?",
      "votes": null
    },
    {
      "id": "3061840",
      "postDate": "12/03/2024 02:50:06",
      "content": "<p><a href=\"https://www.kaggle.com/vigneshwar472\" target=\"_blank\">@vigneshwar472</a> </p>\n<p>We have synthetic data available here: <a href=\"https://cryoetdataportal.czscience.com/depositions/10310\" target=\"_blank\">https://cryoetdataportal.czscience.com/depositions/10310</a> </p>\n<p>If you use copick for fetching the data (<a href=\"https://copick.github.io/copick/examples/tutorials/data_portal/\" target=\"_blank\">docs here</a>), then you can either just add the synthetic dataset to the same copick project (the data distributions have some differences though) or just create an additional copick project so you can handle the datasets differently.</p>\n<p>If you run out of synthetic data, then you can generate your own, which is a bit more involved, but you can use this <a href=\"https://copick.github.io/copick-catalog/polnet/generate-copick-project/0.1.0\" target=\"_blank\">album solution</a> or just <a href=\"https://github.com/copick/copick-catalog/blob/main/solutions/polnet/generate-copick-project/solution.py\" target=\"_blank\">check out the source</a> if that is more comfy.</p>",
      "rawMarkdown": "vigneshwar472 \n\nWe have synthetic data available here: https://cryoetdataportal.czscience.com/depositions/10310 \n\nIf you use copick for fetching the data ([docs here](https://copick.github.io/copick/examples/tutorials/data_portal/)), then you can either just add the synthetic dataset to the same copick project (the data distributions have some differences though) or just create an additional copick project so you can handle the datasets differently.\n\nIf you run out of synthetic data, then you can generate your own, which is a bit more involved, but you can use this [album solution](https://copick.github.io/copick-catalog/polnet/generate-copick-project/0.1.0) or just [check out the source](https://github.com/copick/copick-catalog/blob/main/solutions/polnet/generate-copick-project/solution.py) if that is more comfy.",
      "votes": null
    },
    {
      "id": "3071156",
      "postDate": "12/13/2024 12:14:59",
      "content": "<p><a href=\"https://www.kaggle.com/kharrington\" target=\"_blank\">@kharrington</a> <br>\nI want to download the synthetic data with the same format as the competition data.<br>\nI try below code, how can it convert the download data to the format as the competition data?</p>\n<pre><code> cryoet_data_portal import Client, Dataset\n\n = Client()\n\ndataset = Dataset.get_by_id(, )\ndataset.download_everything()\n</code></pre>",
      "rawMarkdown": "kharrington \nI want to download the synthetic data with the same format as the competition data.\nI try below code, how can it convert the download data to the format as the competition data?\n```\nfrom cryoet_data_portal import Client, Dataset\n\nclient = Client()\n\ndataset = Dataset.get_by_id(client, 10441)\ndataset.download_everything()\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3038900,
      "author_name": "kharrington",
      "author_url": "",
      "post_date": "11/07/2024 13:59:51",
      "content": "<p>I'll make a more detailed post about the limited dataset today. This will include some explanation about <em>why</em> we have provided such little training data.</p>\n<p>There are extra datasets available for training, including synthetic data that matches the material in the test/training images. We have had luck combining the synthetic and experimental data. </p>\n<p>I will make a point of adding more documentation about getting the extra training data. Thank you for pointing all this out!!!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3038909,
          "author_name": "seshurajup",
          "author_url": "",
          "post_date": "11/07/2024 14:06:31",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/kharrington\" target=\"_blank\">@kharrington</a> for the quick response about data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3058943,
          "author_name": "vigneshwar472",
          "author_url": "",
          "post_date": "11/30/2024 06:51:16",
          "content": "<p><a href=\"https://www.kaggle.com/kharrington\" target=\"_blank\">@kharrington</a> </p>\n<p>How can I get extra training data?</p>\n<p>It will be useful. </p>\n<p>Did you create the post or documentation?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3061840,
              "author_name": "kharrington",
              "author_url": "",
              "post_date": "12/03/2024 02:50:06",
              "content": "<p><a href=\"https://www.kaggle.com/vigneshwar472\" target=\"_blank\">@vigneshwar472</a> </p>\n<p>We have synthetic data available here: <a href=\"https://cryoetdataportal.czscience.com/depositions/10310\" target=\"_blank\">https://cryoetdataportal.czscience.com/depositions/10310</a> </p>\n<p>If you use copick for fetching the data (<a href=\"https://copick.github.io/copick/examples/tutorials/data_portal/\" target=\"_blank\">docs here</a>), then you can either just add the synthetic dataset to the same copick project (the data distributions have some differences though) or just create an additional copick project so you can handle the datasets differently.</p>\n<p>If you run out of synthetic data, then you can generate your own, which is a bit more involved, but you can use this <a href=\"https://copick.github.io/copick-catalog/polnet/generate-copick-project/0.1.0\" target=\"_blank\">album solution</a> or just <a href=\"https://github.com/copick/copick-catalog/blob/main/solutions/polnet/generate-copick-project/solution.py\" target=\"_blank\">check out the source</a> if that is more comfy.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3071156,
                  "author_name": "tangtang1999",
                  "author_url": "",
                  "post_date": "12/13/2024 12:14:59",
                  "content": "<p><a href=\"https://www.kaggle.com/kharrington\" target=\"_blank\">@kharrington</a> <br>\nI want to download the synthetic data with the same format as the competition data.<br>\nI try below code, how can it convert the download data to the format as the competition data?</p>\n<pre><code> cryoet_data_portal import Client, Dataset\n\n = Client()\n\ndataset = Dataset.get_by_id(, )\ndataset.download_everything()\n</code></pre>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3039539,
      "author_name": "arunimbasak",
      "author_url": "",
      "post_date": "11/08/2024 06:51:51",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a> ,<br>\nThankyou for providing this valuable code. <br>\nBut I was having some difficulty in understanding some things here. In the heading, you said 27 extra training samples and then you are saying average 1788 objects. I am having confusion in that. It would be helpful if you could please clarify this ?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3038774": "* is it good to use training set as CV ?\n* extra training samples as training ?\n* 500 test set !!!\n\n----\n\n# How to download extra training?\n```python\nfrom cryoet_data_portal import Client, Tomogram\n\nclient = Client()\n\ntomogram = Tomogram.get_by_id(client, 17049) # - 16172 to 16198\ntomogram.download_omezarr() # many times its freeze and too slow. is good host provide these zarr also as external kaggle dataset\n```\n\n----\n\n# Training # of objects - Average 181 objects\n```python\nProcessing Run: TS_5_4\nAnnotating 140 objects ...\n\nProcessing Run: TS_69_2\nAnnotating 143 objects ...\n\nProcessing Run: TS_6_4\nAnnotating 193 objects ...\n\nProcessing Run: TS_6_6\nAnnotating 143 objects ...\n\nProcessing Run: TS_73_6\nAnnotating 217 objects ...\n\nProcessing Run: TS_86_3\nAnnotating 225 objects ...\n\nProcessing Run: TS_99_9\nAnnotating 208 objects ...\n```\n\n# From Github notebook - Average 1788 objects [Dataset - CZII Extra Dataset](https://www.kaggle.com/datasets/seshurajup/czii-extra-dataset/data)\n```python\nProcessing Run: 16172\nAnnotating 1774 objects ...\n\nProcessing Run: 16173\nAnnotating 1786 objects ...\n\nProcessing Run: 16174\nAnnotating 1784 objects ...\n\nProcessing Run: 16175\nAnnotating 1784 objects ...\n\nProcessing Run: 16176\nAnnotating 1781 objects ...\n\nProcessing Run: 16177\nAnnotating 1827 objects ...\n\nProcessing Run: 16178\nAnnotating 1781 objects ...\n\nProcessing Run: 16179\nAnnotating 1811 objects ...\n\nProcessing Run: 16180\nAnnotating 1808 objects ...\n\nProcessing Run: 16181\nAnnotating 1786 objects ...\n\nProcessing Run: 16182\nAnnotating 1777 objects ...\n\nProcessing Run: 16183\nAnnotating 1805 objects ...\n\nProcessing Run: 16184\nAnnotating 1791 objects ...\n\nProcessing Run: 16185\nAnnotating 1806 objects ...\n\nProcessing Run: 16186\nAnnotating 1759 objects ...\n\nProcessing Run: 16187\nAnnotating 1809 objects ...\n\nProcessing Run: 16188\nAnnotating 1770 objects ...\n\nProcessing Run: 16189\nAnnotating 1765 objects ...\n\nProcessing Run: 16190\nAnnotating 1796 objects ...\n\nProcessing Run: 16191\nAnnotating 1801 objects ...\n\nProcessing Run: 16192\nAnnotating 1766 objects ...\n\nProcessing Run: 16193\nAnnotating 1789 objects ...\n\nProcessing Run: 16194\nAnnotating 1799 objects ...\n\nProcessing Run: 16195\nAnnotating 1784 objects ...\n\nProcessing Run: 16196\nAnnotating 1783 objects ...\n\nProcessing Run: 16197\nAnnotating 1824 objects ...\n\nProcessing Run: 16198\nAnnotating 1751 objects ...\n```",
    "3038900": "I'll make a more detailed post about the limited dataset today. This will include some explanation about *why* we have provided such little training data.\n\nThere are extra datasets available for training, including synthetic data that matches the material in the test/training images. We have had luck combining the synthetic and experimental data. \n\nI will make a point of adding more documentation about getting the extra training data. Thank you for pointing all this out!!!",
    "3038909": "Thanks @kharrington for the quick response about data.",
    "3039539": "Hi @seshurajup ,\nThankyou for providing this valuable code. \nBut I was having some difficulty in understanding some things here. In the heading, you said 27 extra training samples and then you are saying average 1788 objects. I am having confusion in that. It would be helpful if you could please clarify this ?",
    "3058943": "kharrington \n\nHow can I get extra training data?\n\nIt will be useful. \n\nDid you create the post or documentation?",
    "3061840": "vigneshwar472 \n\nWe have synthetic data available here: https://cryoetdataportal.czscience.com/depositions/10310 \n\nIf you use copick for fetching the data ([docs here](https://copick.github.io/copick/examples/tutorials/data_portal/)), then you can either just add the synthetic dataset to the same copick project (the data distributions have some differences though) or just create an additional copick project so you can handle the datasets differently.\n\nIf you run out of synthetic data, then you can generate your own, which is a bit more involved, but you can use this [album solution](https://copick.github.io/copick-catalog/polnet/generate-copick-project/0.1.0) or just [check out the source](https://github.com/copick/copick-catalog/blob/main/solutions/polnet/generate-copick-project/solution.py) if that is more comfy.",
    "3071156": "kharrington \nI want to download the synthetic data with the same format as the competition data.\nI try below code, how can it convert the download data to the format as the competition data?\n```\nfrom cryoet_data_portal import Client, Dataset\n\nclient = Client()\n\ndataset = Dataset.get_by_id(client, 10441)\ndataset.download_everything()\n```"
  },
  "source": "meta"
}