{
  "id": 545381,
  "title": "[solved] is it strange that submitted score is almost zero while local cv is in the range of 0.35 to 0.75?",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/545381",
  "author_name": "hengck23",
  "post_date": "2024-11-09T20:00:59.802000",
  "votes": 38,
  "comment_count": 45,
  "views": 0,
  "content": "<p>I trained on 6 volumes and used the remaining one as validation. I could reach around local cv of about 0.35 to 0.75. However, the submitted lb score is always near 0.</p>\n<p>I suspect there is something wrong?</p>\n<p>my approach:</p>\n<ol>\n<li>use only the highest resolution volume array:  [184, 630, 630]</li>\n<li>locate particle object in the array e.g. coord =(z,y,x)</li>\n<li>convert to location = 10*coord and submit</li>\n</ol>\n<p>evaluation script:<br>\n<a href=\"https://www.kaggle.com/code/metric/czi-cryoet-84969\" target=\"_blank\">https://www.kaggle.com/code/metric/czi-cryoet-84969</a></p>\n<p>did you guys get the same funny results?</p>",
  "messages": [
    {
      "id": 3041016,
      "postDate": "2024-11-09T20:00:59.803Z",
      "content": "<p>I trained on 6 volumes and used the remaining one as validation. I could reach around local cv of about 0.35 to 0.75. However, the submitted lb score is always near 0.</p>\n<p>I suspect there is something wrong?</p>\n<p>my approach:</p>\n<ol>\n<li>use only the highest resolution volume array:  [184, 630, 630]</li>\n<li>locate particle object in the array e.g. coord =(z,y,x)</li>\n<li>convert to location = 10*coord and submit</li>\n</ol>\n<p>evaluation script:<br>\n<a href=\"https://www.kaggle.com/code/metric/czi-cryoet-84969\" target=\"_blank\">https://www.kaggle.com/code/metric/czi-cryoet-84969</a></p>\n<p>did you guys get the same funny results?</p>",
      "rawMarkdown": "I trained on 6 volumes and used the remaining one as validation. I could reach around local cv of about 0.35 to 0.75. However, the submitted lb score is always near 0.\n\nI suspect there is something wrong?\n\nmy approach:\n1. use only the highest resolution volume array:  [184, 630, 630]\n2. locate particle object in the array e.g. coord =(z,y,x)\n3. convert to location = 10*coord and submit\n\nevaluation script:\nhttps://www.kaggle.com/code/metric/czi-cryoet-84969\n\ndid you guys get the same funny results?",
      "votes": 38
    },
    {
      "id": 3044890,
      "postDate": "2024-11-13T23:43:38.330Z",
      "content": "<p>Howdy all ( <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> <a href=\"https://www.kaggle.com/davidlist\" target=\"_blank\">@davidlist</a> <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/sacuscreed\" target=\"_blank\">@sacuscreed</a> <a href=\"https://www.kaggle.com/hypocrites\" target=\"_blank\">@hypocrites</a> <a href=\"https://www.kaggle.com/chemdatafarmer\" target=\"_blank\">@chemdatafarmer</a> <a href=\"https://www.kaggle.com/herminw\" target=\"_blank\">@herminw</a> )</p>\n<p>We have identified the bug in evaluation and will update soon.</p>\n<p>Thank you for all of your investigations!<br>\nKyle</p>",
      "rawMarkdown": "Howdy all ( @hengck23 @davidlist @cdeotte @sacuscreed @hypocrites @chemdatafarmer @herminw )\n\nWe have identified the bug in evaluation and will update soon.\n\nThank you for all of your investigations!\nKyle",
      "votes": 19,
      "replies": [
        {
          "id": 3045383,
          "postDate": "2024-11-14T12:50:11.843Z",
          "content": "<p>Will the competition length be increased due to this period of having submissions which where difficult to evaluate?</p>",
          "rawMarkdown": "Will the competition length be increased due to this period of having submissions which where difficult to evaluate?"
        }
      ]
    },
    {
      "id": 3043866,
      "postDate": "2024-11-12T19:00:40.873Z",
      "content": "<p>Why is LB benchmark 0.587 and top Kaggler only have 0.003 after 1 week? This seems like a bug in Kaggle's scoring system. Has anyone contacted Kaggle about this?</p>",
      "rawMarkdown": "Why is LB benchmark 0.587 and top Kaggler only have 0.003 after 1 week? This seems like a bug in Kaggle's scoring system. Has anyone contacted Kaggle about this?",
      "votes": 14
    },
    {
      "id": 3041050,
      "postDate": "2024-11-09T21:31:32.040Z",
      "content": "<p>It does look like something is weird since the leaderboard has a benchmark at 0.587 but the best submission is only 0.002</p>",
      "rawMarkdown": "It does look like something is weird since the leaderboard has a benchmark at 0.587 but the best submission is only 0.002",
      "votes": 9
    },
    {
      "id": 3044346,
      "postDate": "2024-11-13T09:38:59.953Z",
      "content": "<p>this is something interesting: local cv score for randomly generated 100 points:</p>\n<pre><code>['TS_5_4', 'TS_6_4', 'TS_6_6', 'TS_69_2', 'TS_73_6', 'TS_86_3', 'TS_99_9']\nlb_score 0.0026500562150282825\n         particle_type    P    T  hit  miss   fp  precision    recall   f-beta4  weight\n0         apo-ferritin      0.000000  0.000000  0.000000       1\n1         beta-amylase        0.000000  0.000000  0.000000       0\n2   beta-galactosidase      0.000000  0.000000  0.000000       2\n3             ribosome      0.005714  0.012085  0.011341       1\n4        thyroglobulin      0.001429  0.003984  0.003605       2\n5  virus-like-particle      0.000000  0.000000  0.000000       1\n</code></pre>\n<pre><code>submit_df=\n id  SIMPLE_ID:\n     n  PARTICLE_NAME:\n        (id,n)\n        D,H,W = , ,  \n        xyz=np(,,(,)) \n        xyz = xyz*]+]\n        xyz = xyz* \n        submit_df(\n            pd({: id, : n, : xyz, : xyz, : xyz})\n        ) \nsubmit_df = pd(submit_df)\nsubmit_df(loc=, column=, value=np((submit_df)))\n\n</code></pre>",
      "rawMarkdown": "this is something interesting: local cv score for randomly generated 100 points:\n\n```\n\n['TS_5_4', 'TS_6_4', 'TS_6_6', 'TS_69_2', 'TS_73_6', 'TS_86_3', 'TS_99_9']\nlb_score 0.0026500562150282825\n         particle_type    P    T  hit  miss   fp  precision    recall   f-beta4  weight\n0         apo-ferritin  700  375    0   375  700   0.000000  0.000000  0.000000       1\n1         beta-amylase  700   87    0    87  700   0.000000  0.000000  0.000000       0\n2   beta-galactosidase  700  112    0   112  700   0.000000  0.000000  0.000000       2\n3             ribosome  700  331    4   327  696   0.005714  0.012085  0.011341       1\n4        thyroglobulin  700  251    1   250  699   0.001429  0.003984  0.003605       2\n5  virus-like-particle  700  113    0   113  700   0.000000  0.000000  0.000000       1\n\n```\n\n\n```\nsubmit_df=[]\nfor id in SIMPLE_ID:\n    for n in PARTICLE_NAME:\n        print(id,n)\n        D,H,W = 184, 630, 630 \n        xyz=np.random.uniform(0,1,(100,3)) \n        xyz = xyz*[[W,H,0.5*D]]+[[0,0,0.25*D]]\n        xyz = xyz*10 \n        submit_df.append(\n            pd.DataFrame({'experiment': id, 'particle_type': n, 'x': xyz[:, 0], 'y': xyz[:, 1], 'z': xyz[:, 2]})\n        ) \nsubmit_df = pd.concat(submit_df)\nsubmit_df.insert(loc=0, column='id', value=np.arange(len(submit_df)))\nprint(submit_df)\n\n```",
      "votes": 3,
      "replies": [
        {
          "id": 3044358,
          "postDate": "2024-11-13T10:04:20.650Z",
          "content": "<p>Nice!  Looks like I've got a shot now!  …Just gotta optimize number of random points for f-score…and done!  😀</p>",
          "rawMarkdown": "Nice!  Looks like I've got a shot now!  ...Just gotta optimize number of random points for f-score...and done!  😀",
          "votes": 2
        },
        {
          "id": 3044360,
          "postDate": "2024-11-13T10:12:00.583Z",
          "content": "<p>further:</p>\n<pre><code> of random points per tomograph, local cv, public score:\n, ., .\n, .\n, .\n,  .\n</code></pre>\n<p><a href=\"https://www.kaggle.com/code/hengck23/random-submission\" target=\"_blank\">https://www.kaggle.com/code/hengck23/random-submission</a></p>\n<p>i probably can conclude that the train and test labels are different.<br>\nAnd the difference either come from bug or the data itself.</p>\n<p>(I suspect there is wrong ordering of run id (or particle type) in the ground  truth)</p>\n<hr>\n<pre><code>:.\n    =np.random.uniform(,,(,)) \n     = xyz*[[W,H,D]]\n     = xyz**.\n\n\n:.\n    =np.random.uniform(,,(,)) \n     = xyz*[[W,H,D]]\n     = xyz**\n</code></pre>",
          "rawMarkdown": "further:\n```\nnum of random points per tomograph, local cv, public score:\n800, 0.007004517, 0.003\n500, 0.00699\n200, 0.004255\n100,  0.001735\n\n```\nhttps://www.kaggle.com/code/hengck23/random-submission\n\ni probably can conclude that the train and test labels are different.\nAnd the difference either come from bug or the data itself.\n\n(I suspect there is wrong ordering of run id (or particle type) in the ground  truth)\n\n\n----\n```\nlb:0.003\n    xyz=np.random.uniform(0,1,(800,3)) \n    xyz = xyz*[[W,H,D]]\n    xyz = xyz*10*0.5\n\n\nlb:0.000\n    xyz=np.random.uniform(0,1,(800,3)) \n    xyz = xyz*[[W,H,D]]\n    xyz = xyz*10*2\n```",
          "votes": 3,
          "replies": [
            {
              "id": 3044368,
              "postDate": "2024-11-13T10:18:12.007Z",
              "content": "<p>Looks like 2000 gets you to the top of the current leaderboard.  Pretty fantasic analysis, BTW!</p>",
              "rawMarkdown": "Looks like 2000 gets you to the top of the current leaderboard.  Pretty fantasic analysis, BTW!",
              "votes": 1
            },
            {
              "id": 3044479,
              "postDate": "2024-11-13T12:53:42.043Z",
              "content": "<p>Have anyone tryed with no transformation at all? Raw xyz predictions. I mean, for training sure. You need to apply the scale in order to make sense. But may be in submission values are raw or transformerd after automatically.</p>",
              "rawMarkdown": "Have anyone tryed with no transformation at all? Raw xyz predictions. I mean, for training sure. You need to apply the scale in order to make sense. But may be in submission values are raw or transformerd after automatically."
            },
            {
              "id": 3044598,
              "postDate": "2024-11-13T15:24:34.127Z",
              "content": "<p>Interesting finding</p>",
              "rawMarkdown": "Interesting finding"
            }
          ]
        }
      ]
    },
    {
      "id": 3042243,
      "postDate": "2024-11-11T10:57:57.943Z",
      "content": "<p><a href=\"https://www.kaggle.com/kharrington\" target=\"_blank\">@kharrington</a> what do you think about this? Seems there is something weird with submission evaluation.</p>",
      "rawMarkdown": "@kharrington what do you think about this? Seems there is something weird with submission evaluation.",
      "votes": 3,
      "replies": [
        {
          "id": 3042613,
          "postDate": "2024-11-11T16:10:43.790Z",
          "content": "<p>I agree, this is strange. Hopefully we can get some clarification.</p>",
          "rawMarkdown": "I agree, this is strange. Hopefully we can get some clarification.",
          "votes": 1
        }
      ]
    },
    {
      "id": 3041167,
      "postDate": "2024-11-10T02:16:01.120Z",
      "content": "<p>The sample submission notebook using blob detector: indeed zero for local lb</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fbaf2c5967cb015e9ad1417dd38f93857%2FSelection_657.png?generation=1731204927157155&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "The sample submission notebook using blob detector: indeed zero for local lb\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fbaf2c5967cb015e9ad1417dd38f93857%2FSelection_657.png?generation=1731204927157155&alt=media)",
      "votes": 4,
      "replies": [
        {
          "id": 3041176,
          "postDate": "2024-11-10T02:29:40.247Z",
          "content": "<p>I wonder if it's as simple as applying a constant correction to the output, although I'm not sure I'm seeing consistency in the offset.</p>",
          "rawMarkdown": "I wonder if it's as simple as applying a constant correction to the output, although I'm not sure I'm seeing consistency in the offset.",
          "votes": 2,
          "replies": [
            {
              "id": 3041342,
              "postDate": "2024-11-10T07:43:04.827Z",
              "content": "<p>even if you correct the offset error, your lb score is still zero.<br>\nthe results of the simple blob detector are too bad.</p>",
              "rawMarkdown": "even if you correct the offset error, your lb score is still zero.\nthe results of the simple blob detector are too bad.",
              "votes": 1
            },
            {
              "id": 3041838,
              "postDate": "2024-11-10T23:34:00.377Z",
              "content": "<p>I was thinking you meant there was an offset between the way the grid is described and how it is evaluated for the leaderboard, but of course, that's not something you'd be able to display.</p>",
              "rawMarkdown": "I was thinking you meant there was an offset between the way the grid is described and how it is evaluated for the leaderboard, but of course, that's not something you'd be able to display.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3041126,
      "postDate": "2024-11-10T01:12:06.943Z",
      "content": "<p>Similarly, cv0.6 lb0.002, and I believe that the benchmark on the leaderboard was not evaluated through Kaggle's notebook, it is possible that only the CSV was uploaded. I have also tried DeepFindET for benchmarks, and the inference time alone exceeded 12 hours,  The time of coordinate conversion is not included</p>",
      "rawMarkdown": "Similarly, cv0.6 lb0.002, and I believe that the benchmark on the leaderboard was not evaluated through Kaggle's notebook, it is possible that only the CSV was uploaded. I have also tried DeepFindET for benchmarks, and the inference time alone exceeded 12 hours,  The time of coordinate conversion is not included",
      "votes": 4,
      "replies": [
        {
          "id": 3041166,
          "postDate": "2024-11-10T02:15:05.947Z",
          "content": "<p>i think there is some bug in the evaluation process. (e.g. wrong truth csv file; or x,y,z coord get mixed up; or wrong scaling of 10)</p>\n<p>Now that DeepFindET gets 0.5 in the benchmark. It is pretrained with synthetic + fine-tune on kaggle train data.<br>\nso if we remove the pretrained step, DeepFindET score would then be at least 0.25 (very worst case)<br>\nDeepFindET is just a 3d unet, so I think any 3d unet should score at least 0.15 to  0.20.</p>",
          "rawMarkdown": "i think there is some bug in the evaluation process. (e.g. wrong truth csv file; or x,y,z coord get mixed up; or wrong scaling of 10)\n\nNow that DeepFindET gets 0.5 in the benchmark. It is pretrained with synthetic + fine-tune on kaggle train data.\nso if we remove the pretrained step, DeepFindET score would then be at least 0.25 (very worst case)\nDeepFindET is just a 3d unet, so I think any 3d unet should score at least 0.15 to  0.20.",
          "replies": [
            {
              "id": 3041357,
              "postDate": "2024-11-10T08:02:34.800Z",
              "content": "<p>I doubt this is it, but the data arrays are definitely z, y, x as far as order of indices.</p>",
              "rawMarkdown": "I doubt this is it, but the data arrays are definitely z, y, x as far as order of indices."
            },
            {
              "id": 3042665,
              "postDate": "2024-11-11T16:54:09.907Z",
              "content": "<p>baseline model: 0.001<br>\nflip z: lb=0<br>\ntranspose x,y: lb=0<br>\nremove 10 scale: lb=0<br>\n2 class (background versus particle): 0.003<br>\nraise error if average num of detected particle is less than 10:  in submission</p>\n<p>i think i know where we are getting to …</p>",
              "rawMarkdown": "baseline model: 0.001\nflip z: lb=0\ntranspose x,y: lb=0\nremove 10 scale: lb=0\n2 class (background versus particle): 0.003\nraise error if average num of detected particle is less than 10:  in submission\n\ni think i know where we are getting to ...",
              "votes": 1
            }
          ]
        },
        {
          "id": 3041170,
          "postDate": "2024-11-10T02:19:52.273Z",
          "content": "<p>\" time alone exceeded 12 hours, \"<br>\nyou can just run DeepFindET e.g. for 50% of the test sample (then we expect a lower lb score, but not zero).</p>\n<p>But I think there is issue with the evaluation server and your local cv0.60 is correct</p>",
          "rawMarkdown": "\" time alone exceeded 12 hours, \"\nyou can just run DeepFindET e.g. for 50% of the test sample (then we expect a lower lb score, but not zero).\n\nBut I think there is issue with the evaluation server and your local cv0.60 is correct",
          "votes": 1
        }
      ]
    },
    {
      "id": 3042196,
      "postDate": "2024-11-11T09:56:42.253Z",
      "content": "<p>the context of conversion is important.  (omitted because it is too long).<br>\nbasically I ask him lots of question of SNR, etc … those question that I do not understand from reading papers before asking him the kaggle question.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F47093115e66cb51f4f31af7b2d65e2f8%2FSelection_665.png?generation=1731318874920460&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F9f770d39b5398a85e582c52105306ceb%2FSelection_666.png?generation=1731319233851404&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "the context of conversion is important.  (omitted because it is too long).\nbasically I ask him lots of question of SNR, etc ... those question that I do not understand from reading papers before asking him the kaggle question.\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F47093115e66cb51f4f31af7b2d65e2f8%2FSelection_665.png?generation=1731318874920460&alt=media)\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F9f770d39b5398a85e582c52105306ceb%2FSelection_666.png?generation=1731319233851404&alt=media)",
      "votes": 1,
      "replies": [
        {
          "id": 3043065,
          "postDate": "2024-11-12T04:23:43.757Z",
          "content": "<p>what is \"SNR\"? explain, please.</p>",
          "rawMarkdown": "what is \"SNR\"? explain, please.",
          "replies": [
            {
              "id": 3043136,
              "postDate": "2024-11-12T05:42:56.043Z",
              "content": "<p>signal to noise ratio</p>",
              "rawMarkdown": "signal to noise ratio",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 3041855,
      "postDate": "2024-11-11T00:11:43.093Z",
      "content": "<p>does the public test data contins simulated data only?<br>\n(and private contain both simulated and real data)</p>\n<p>we are given 7 real train data (and ask to use the  simulated data from the portal ourselves)<br>\nis this the explanation for zero score????</p>\n<hr>\n<p>please note the information from the dataset paper:</p>\n<p>Quotes: </p>\n<ol>\n<li><p>Dataset split: We curated a selection of <strong>492</strong> good-quality tomograms (Extended Data Fig. 7)<br>\nfrom the phantom dataset and divided them into three subsets: training, testing, and validation.</p></li>\n<li><p>All training data described here (7 tomograms) as well as additional published training data are<br>\navailable on the CZ CryoET Data Portal (CZCDP16, cryoetdataportal.czscience.com) under<br>\ndataset IDs 10440 (experimental data) and 10441 (PolNet43-simulations).</p></li>\n</ol>\n<hr>\n<p><a href=\"https://cryoetdataportal.czscience.com/browse-data/datasets\" target=\"_blank\">https://cryoetdataportal.czscience.com/browse-data/datasets</a><br>\nDataset ID: DS-10440 (7 runs)</p>\n<ul>\n<li>The data was acquired on a Krios G4 using a Falcon 4i detector and SelctrisX energy filter. Tomograms were reconstructed using AreTomo3 v1.0.23 and post-processed using different methods. Each run provides raw, denoised, missing wedge corrected and ctf corrected tomograms. </li>\n</ul>\n<p>Dataset ID: DS-10441 (27 runs)</p>\n<ul>\n<li>simulated tiltseries, tomograms and ground truth annotations for the purpose of training object identification algorithms in the context of the CryoET Object Identification Challenge. The data was simulated using polnet, and contains point labels</li>\n</ul>\n<hr>\n<p>now in  <a href=\"https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks/blob/main/DeepFindET/train.ipynb:\" target=\"_blank\">https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks/blob/main/DeepFindET/train.ipynb:</a></p>\n<p>Quotes: </p>\n<ul>\n<li>we demonstrate how to utilize this infrastructure to predict the 3D coordinates of six proteins of varying sizes, provided by the CryoET Dataportal (Dataset ID: 10439).  …</li>\n<li>copick config</li>\n</ul>\n<pre><code> : ,\n : ,\n   \n: ,\n            : ,\n            : 8,\n</code></pre>\n<p>is  CryoET Dataportal (Dataset ID: 10439) a miske? shouldn't it be 10441????</p>",
      "rawMarkdown": "does the public test data contins simulated data only?\n(and private contain both simulated and real data)\n\nwe are given 7 real train data (and ask to use the  simulated data from the portal ourselves)\nis this the explanation for zero score????\n\n---\n\nplease note the information from the dataset paper:\n\nQuotes: \n\n1. Dataset split: We curated a selection of **492** good-quality tomograms (Extended Data Fig. 7)\nfrom the phantom dataset and divided them into three subsets: training, testing, and validation.\n\n2. All training data described here (7 tomograms) as well as additional published training data are\navailable on the CZ CryoET Data Portal (CZCDP16, cryoetdataportal.czscience.com) under\ndataset IDs 10440 (experimental data) and 10441 (PolNet43-simulations).\n\n---\n\nhttps://cryoetdataportal.czscience.com/browse-data/datasets\nDataset ID: DS-10440 (7 runs)\n- The data was acquired on a Krios G4 using a Falcon 4i detector and SelctrisX energy filter. Tomograms were reconstructed using AreTomo3 v1.0.23 and post-processed using different methods. Each run provides raw, denoised, missing wedge corrected and ctf corrected tomograms. \n\n\nDataset ID: DS-10441 (27 runs)\n- simulated tiltseries, tomograms and ground truth annotations for the purpose of training object identification algorithms in the context of the CryoET Object Identification Challenge. The data was simulated using polnet, and contains point labels\n\n\n---\n\nnow in  https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks/blob/main/DeepFindET/train.ipynb:\n\n\nQuotes: \n\n-  we demonstrate how to utilize this infrastructure to predict the 3D coordinates of six proteins of varying sizes, provided by the CryoET Dataportal (Dataset ID: 10439).  ...\n-  copick config\n```\n \"name\": \"apo-ferritin\",\n \"name\": \"beta-amylase\",\n...   \n\"name\": \"membrane\",\n            \"is_particle\": false,\n            \"label\": 8,\n\n```\n\nis  CryoET Dataportal (Dataset ID: 10439) a miske? shouldn't it be 10441????\n",
      "votes": 1,
      "replies": [
        {
          "id": 3042025,
          "postDate": "2024-11-11T05:02:30.577Z",
          "content": "<p>Looking at the <strong>Annotated Objects</strong> given in <a href=\"https://cryoetdataportal.czscience.com/datasets/10439\" target=\"_blank\">10439</a> and <a href=\"https://cryoetdataportal.czscience.com/datasets/10441\" target=\"_blank\">10441</a>, you seem to be right, it should be 10441.</p>",
          "rawMarkdown": "Looking at the **Annotated Objects** given in [10439](https://cryoetdataportal.czscience.com/datasets/10439) and [10441](https://cryoetdataportal.czscience.com/datasets/10441), you seem to be right, it should be 10441.",
          "replies": [
            {
              "id": 3043622,
              "postDate": "2024-11-12T14:41:25.003Z",
              "content": "<p>Ah, sorry about the dataset ID mix up there. When making the notebooks we used a different dataset to ensure that no data/info about the competition leaked, so we had an alternative synthetic dataset. You are correct 10441 is the synthetic dataset to use.</p>",
              "rawMarkdown": "Ah, sorry about the dataset ID mix up there. When making the notebooks we used a different dataset to ensure that no data/info about the competition leaked, so we had an alternative synthetic dataset. You are correct 10441 is the synthetic dataset to use.",
              "votes": 3
            }
          ]
        },
        {
          "id": 3043626,
          "postDate": "2024-11-12T14:48:35.063Z",
          "content": "<p>to clarify this:</p>\n<pre><code>does the  test  contins simulated  ?\n(and  contain both simulated and  )\n\nwe are given   train  (and ask to  the simulated  from the portal ourselves)\nis this the explanation for zero score????\n</code></pre>\n<p>Public and private test data <em>only</em> contain real data. Simulated data is only provided to supplement training data.</p>",
          "rawMarkdown": "to clarify this:\n\n```\ndoes the public test data contins simulated data only?\n(and private contain both simulated and real data)\n\nwe are given 7 real train data (and ask to use the simulated data from the portal ourselves)\nis this the explanation for zero score????\n```\n\nPublic and private test data *only* contain real data. Simulated data is only provided to supplement training data.",
          "votes": 5
        }
      ]
    },
    {
      "id": 3041171,
      "postDate": "2024-11-10T02:22:45.910Z",
      "content": "<p>Interesting.  I submitted a ribosome only model that should have found at least something, but didn't even move up the leaderboard with a better 0.000 than before. 🤔</p>",
      "rawMarkdown": "Interesting.  I submitted a ribosome only model that should have found at least something, but didn't even move up the leaderboard with a better 0.000 than before. 🤔",
      "votes": 1,
      "replies": [
        {
          "id": 3041179,
          "postDate": "2024-11-10T02:41:59.327Z",
          "content": "<p>Oh, wait!  I <strong>did</strong> get a better 0.000.  Nice!  I'm still in the money!  😀</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10704200%2F467cb35e09122bd7226b49348b64e64c%2Fkaggle.jpg?generation=1731206452204999&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Oh, wait!  I **did** get a better 0.000.  Nice!  I'm still in the money!  😀\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10704200%2F467cb35e09122bd7226b49348b64e64c%2Fkaggle.jpg?generation=1731206452204999&alt=media)",
          "votes": 1
        }
      ]
    },
    {
      "id": 3041082,
      "postDate": "2024-11-09T22:48:35.997Z",
      "content": "<p>My disparity is far from yours, but I also experienced it nonetheless: local CV of ~0.013 -&gt; LB of ~0.000.</p>",
      "rawMarkdown": "My disparity is far from yours, but I also experienced it nonetheless: local CV of ~0.013 -> LB of ~0.000.",
      "votes": 1
    },
    {
      "id": 3041364,
      "postDate": "2024-11-10T08:20:04.060Z",
      "content": "<p>Hi, Heng</p>\n<p>Your scaling \"10*coord and submit\" is a bit off.<br>\nCorrect scales are \"scale\": [10.012444196428572, 10.012444196428572, 10.012444537618887]<br>\nThe difference is tiny though</p>",
      "rawMarkdown": "Hi, Heng\n\nYour scaling \"10*coord and submit\" is a bit off.\nCorrect scales are \"scale\": [10.012444196428572, 10.012444196428572, 10.012444537618887]\nThe difference is tiny though",
      "votes": 2,
      "replies": [
        {
          "id": 3041392,
          "postDate": "2024-11-10T09:16:24.163Z",
          "content": "<p><a href=\"https://www.kaggle.com/sakvaua\" target=\"_blank\">@sakvaua</a> Where are you getting that 10.012444196428572 from?</p>",
          "rawMarkdown": "@sakvaua Where are you getting that 10.012444196428572 from?",
          "replies": [
            {
              "id": 3041415,
              "postDate": "2024-11-10T10:14:17.840Z",
              "content": "<p>Host confirmed in <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/544895#3040071\" target=\"_blank\">here</a>, based on zarrs.</p>",
              "rawMarkdown": "Host confirmed in [here](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/544895#3040071), based on zarrs.",
              "votes": 1
            },
            {
              "id": 3041416,
              "postDate": "2024-11-10T10:15:01.363Z",
              "content": "<p>Zarr attributes show this:<br>\n{\"multiscales\": [{\"axes\": [{\"name\": \"z\", \"type\": \"space\", \"unit\": \"angstrom\"}, {\"name\": \"y\", \"type\": \"space\", \"unit\": \"angstrom\"}, {\"name\": \"x\", \"type\": \"space\", \"unit\": \"angstrom\"}], \"datasets\": [{\"coordinateTransformations\": [{\"scale\": [10.012444196428572, 10.012444196428572, 10.012444537618887], \"type\": \"scale\"}], \"path\": \"0\"}, {\"coordinateTransformations\": [{\"scale\": [20.024888392857143, 20.024888392857143, 20.024889075237773], \"type\": \"scale\"}], \"path\": \"1\"}, {\"coordinateTransformations\": [{\"scale\": [40.049776785714286, 40.049776785714286, 40.04977815047555], \"type\": \"scale\"}], \"path\": \"2\"}], \"metadata\": {}, \"name\": \"/\", \"version\": \"0.4\"}]}</p>\n<p>scales for multiscale zars are 10.012, 20.024 and 40.049<br>\nI hope I'm not misinterpreting this</p>",
              "rawMarkdown": "Zarr attributes show this:\n{\"multiscales\": [{\"axes\": [{\"name\": \"z\", \"type\": \"space\", \"unit\": \"angstrom\"}, {\"name\": \"y\", \"type\": \"space\", \"unit\": \"angstrom\"}, {\"name\": \"x\", \"type\": \"space\", \"unit\": \"angstrom\"}], \"datasets\": [{\"coordinateTransformations\": [{\"scale\": [10.012444196428572, 10.012444196428572, 10.012444537618887], \"type\": \"scale\"}], \"path\": \"0\"}, {\"coordinateTransformations\": [{\"scale\": [20.024888392857143, 20.024888392857143, 20.024889075237773], \"type\": \"scale\"}], \"path\": \"1\"}, {\"coordinateTransformations\": [{\"scale\": [40.049776785714286, 40.049776785714286, 40.04977815047555], \"type\": \"scale\"}], \"path\": \"2\"}], \"metadata\": {}, \"name\": \"/\", \"version\": \"0.4\"}]}\n\nscales for multiscale zars are 10.012, 20.024 and 40.049\nI hope I'm not misinterpreting this",
              "votes": 4
            },
            {
              "id": 3041448,
              "postDate": "2024-11-10T11:21:17.323Z",
              "content": "<p>Thanks!  That's pretty compelling.  😀</p>",
              "rawMarkdown": "Thanks!  That's pretty compelling.  😀"
            },
            {
              "id": 3047078,
              "postDate": "2024-11-16T08:15:35.667Z",
              "content": "<p>Actually, the scaling does start to matter for the apo-ferritin which only has a half-radius of 3 pixels.  You can visually see them off center if you don't do the scaling right.</p>",
              "rawMarkdown": "Actually, the scaling does start to matter for the apo-ferritin which only has a half-radius of 3 pixels.  You can visually see them off center if you don't do the scaling right.",
              "votes": 1
            }
          ]
        },
        {
          "id": 3041845,
          "postDate": "2024-11-10T23:53:07.613Z",
          "content": "<p>7.56 units worst case (0.012 x 630).  Tiny, but might be the difference for some of the smaller particles in the lower right corner.</p>",
          "rawMarkdown": "7.56 units worst case (0.012 x 630).  Tiny, but might be the difference for some of the smaller particles in the lower right corner.",
          "votes": 1
        },
        {
          "id": 3047079,
          "postDate": "2024-11-16T08:17:33.393Z",
          "content": "<p>Looks like it actually starts to matter for the apo-ferratin with a half radius of only 3 pixels.  You can visually see them off center without the scaling.</p>",
          "rawMarkdown": "Looks like it actually starts to matter for the apo-ferratin with a half radius of only 3 pixels.  You can visually see them off center without the scaling.",
          "votes": 1
        }
      ]
    },
    {
      "id": 3078340,
      "postDate": "2024-12-22T07:23:08.717Z",
      "content": "<p>Thank you for sharing, it has been very helpful to me</p>",
      "rawMarkdown": "Thank you for sharing, it has been very helpful to me"
    },
    {
      "id": 3044900,
      "postDate": "2024-11-14T00:20:28.630Z",
      "content": "<p>Pure speculation, but I keep getting exceptions/scoring errors when I submit despite the fact that everything works perfectly fine on the provided test/training data.  Wonder if maybe everyone is seeing that same issue but for most it causes only a few files to be processed instead of the exception I'm seeing.</p>",
      "rawMarkdown": "Pure speculation, but I keep getting exceptions/scoring errors when I submit despite the fact that everything works perfectly fine on the provided test/training data.  Wonder if maybe everyone is seeing that same issue but for most it causes only a few files to be processed instead of the exception I'm seeing.",
      "replies": [
        {
          "id": 3044901,
          "postDate": "2024-11-14T00:26:57.357Z",
          "content": "<p>Scratch that.  I was iterating the test files but loading from the training directory.  😐</p>",
          "rawMarkdown": "Scratch that.  I was iterating the test files but loading from the training directory.  😐"
        }
      ]
    },
    {
      "id": 3044703,
      "postDate": "2024-11-13T18:06:13.700Z",
      "content": "<p>It would be nice if <a href=\"https://www.kaggle.com/kharrington\" target=\"_blank\">@kharrington</a> and/or Kaggle staff could confirm here that the organizers recognize something weird is going on and they are looking into it.</p>",
      "rawMarkdown": "It would be nice if @kharrington and/or Kaggle staff could confirm here that the organizers recognize something weird is going on and they are looking into it."
    },
    {
      "id": 3043765,
      "postDate": "2024-11-12T16:57:47.167Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3041217,
      "postDate": "2024-11-10T03:22:47.047Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3044890,
      "author_name": "Kyle Harrington",
      "author_url": "",
      "post_date": "2024-11-13T23:43:38.330000",
      "content": "<p>Howdy all ( <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> <a href=\"https://www.kaggle.com/davidlist\" target=\"_blank\">@davidlist</a> <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/sacuscreed\" target=\"_blank\">@sacuscreed</a> <a href=\"https://www.kaggle.com/hypocrites\" target=\"_blank\">@hypocrites</a> <a href=\"https://www.kaggle.com/chemdatafarmer\" target=\"_blank\">@chemdatafarmer</a> <a href=\"https://www.kaggle.com/herminw\" target=\"_blank\">@herminw</a> )</p>\n<p>We have identified the bug in evaluation and will update soon.</p>\n<p>Thank you for all of your investigations!<br>\nKyle</p>",
      "votes": 19,
      "replies": [
        {
          "id": 3045383,
          "author_name": "stefanoclss",
          "author_url": "",
          "post_date": "2024-11-14T12:50:11.843000",
          "content": "<p>Will the competition length be increased due to this period of having submissions which where difficult to evaluate?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3043866,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2024-11-12T19:00:40.873000",
      "content": "<p>Why is LB benchmark 0.587 and top Kaggler only have 0.003 after 1 week? This seems like a bug in Kaggle's scoring system. Has anyone contacted Kaggle about this?</p>",
      "votes": 14,
      "replies": []
    },
    {
      "id": 3041050,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2024-11-09T21:31:32.040000",
      "content": "<p>It does look like something is weird since the leaderboard has a benchmark at 0.587 but the best submission is only 0.002</p>",
      "votes": 9,
      "replies": []
    },
    {
      "id": 3044346,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-11-13T09:38:59.953000",
      "content": "<p>this is something interesting: local cv score for randomly generated 100 points:</p>\n<pre><code>['TS_5_4', 'TS_6_4', 'TS_6_6', 'TS_69_2', 'TS_73_6', 'TS_86_3', 'TS_99_9']\nlb_score 0.0026500562150282825\n         particle_type    P    T  hit  miss   fp  precision    recall   f-beta4  weight\n0         apo-ferritin      0.000000  0.000000  0.000000       1\n1         beta-amylase        0.000000  0.000000  0.000000       0\n2   beta-galactosidase      0.000000  0.000000  0.000000       2\n3             ribosome      0.005714  0.012085  0.011341       1\n4        thyroglobulin      0.001429  0.003984  0.003605       2\n5  virus-like-particle      0.000000  0.000000  0.000000       1\n</code></pre>\n<pre><code>submit_df=\n id  SIMPLE_ID:\n     n  PARTICLE_NAME:\n        (id,n)\n        D,H,W = , ,  \n        xyz=np(,,(,)) \n        xyz = xyz*]+]\n        xyz = xyz* \n        submit_df(\n            pd({: id, : n, : xyz, : xyz, : xyz})\n        ) \nsubmit_df = pd(submit_df)\nsubmit_df(loc=, column=, value=np((submit_df)))\n\n</code></pre>",
      "votes": 3,
      "replies": [
        {
          "id": 3044358,
          "author_name": "David List",
          "author_url": "",
          "post_date": "2024-11-13T10:04:20.650000",
          "content": "<p>Nice!  Looks like I've got a shot now!  …Just gotta optimize number of random points for f-score…and done!  😀</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 3044360,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2024-11-13T10:12:00.583000",
          "content": "<p>further:</p>\n<pre><code> of random points per tomograph, local cv, public score:\n, ., .\n, .\n, .\n,  .\n</code></pre>\n<p><a href=\"https://www.kaggle.com/code/hengck23/random-submission\" target=\"_blank\">https://www.kaggle.com/code/hengck23/random-submission</a></p>\n<p>i probably can conclude that the train and test labels are different.<br>\nAnd the difference either come from bug or the data itself.</p>\n<p>(I suspect there is wrong ordering of run id (or particle type) in the ground  truth)</p>\n<hr>\n<pre><code>:.\n    =np.random.uniform(,,(,)) \n     = xyz*[[W,H,D]]\n     = xyz**.\n\n\n:.\n    =np.random.uniform(,,(,)) \n     = xyz*[[W,H,D]]\n     = xyz**\n</code></pre>",
          "votes": 3,
          "replies": [
            {
              "id": 3044368,
              "author_name": "David List",
              "author_url": "",
              "post_date": "2024-11-13T10:18:12.007000",
              "content": "<p>Looks like 2000 gets you to the top of the current leaderboard.  Pretty fantasic analysis, BTW!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3044479,
              "author_name": "Ángel Jacinto Sánchez Ruiz",
              "author_url": "",
              "post_date": "2024-11-13T12:53:42.043000",
              "content": "<p>Have anyone tryed with no transformation at all? Raw xyz predictions. I mean, for training sure. You need to apply the scale in order to make sense. But may be in submission values are raw or transformerd after automatically.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3044598,
              "author_name": "Zhuoqun Li",
              "author_url": "",
              "post_date": "2024-11-13T15:24:34.127000",
              "content": "<p>Interesting finding</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3042243,
      "author_name": "anthony",
      "author_url": "",
      "post_date": "2024-11-11T10:57:57.943000",
      "content": "<p><a href=\"https://www.kaggle.com/kharrington\" target=\"_blank\">@kharrington</a> what do you think about this? Seems there is something weird with submission evaluation.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3042613,
          "author_name": "chemdatafarmer",
          "author_url": "",
          "post_date": "2024-11-11T16:10:43.790000",
          "content": "<p>I agree, this is strange. Hopefully we can get some clarification.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3041167,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-11-10T02:16:01.120000",
      "content": "<p>The sample submission notebook using blob detector: indeed zero for local lb</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fbaf2c5967cb015e9ad1417dd38f93857%2FSelection_657.png?generation=1731204927157155&amp;alt=media\" alt=\"\"></p>",
      "votes": 4,
      "replies": [
        {
          "id": 3041176,
          "author_name": "David List",
          "author_url": "",
          "post_date": "2024-11-10T02:29:40.247000",
          "content": "<p>I wonder if it's as simple as applying a constant correction to the output, although I'm not sure I'm seeing consistency in the offset.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3041342,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-11-10T07:43:04.827000",
              "content": "<p>even if you correct the offset error, your lb score is still zero.<br>\nthe results of the simple blob detector are too bad.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3041838,
              "author_name": "David List",
              "author_url": "",
              "post_date": "2024-11-10T23:34:00.377000",
              "content": "<p>I was thinking you meant there was an offset between the way the grid is described and how it is evaluated for the leaderboard, but of course, that's not something you'd be able to display.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3041126,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-11-10T01:12:06.943000",
      "content": "<p>Similarly, cv0.6 lb0.002, and I believe that the benchmark on the leaderboard was not evaluated through Kaggle's notebook, it is possible that only the CSV was uploaded. I have also tried DeepFindET for benchmarks, and the inference time alone exceeded 12 hours,  The time of coordinate conversion is not included</p>",
      "votes": 4,
      "replies": [
        {
          "id": 3041166,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2024-11-10T02:15:05.947000",
          "content": "<p>i think there is some bug in the evaluation process. (e.g. wrong truth csv file; or x,y,z coord get mixed up; or wrong scaling of 10)</p>\n<p>Now that DeepFindET gets 0.5 in the benchmark. It is pretrained with synthetic + fine-tune on kaggle train data.<br>\nso if we remove the pretrained step, DeepFindET score would then be at least 0.25 (very worst case)<br>\nDeepFindET is just a 3d unet, so I think any 3d unet should score at least 0.15 to  0.20.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3041357,
              "author_name": "David List",
              "author_url": "",
              "post_date": "2024-11-10T08:02:34.800000",
              "content": "<p>I doubt this is it, but the data arrays are definitely z, y, x as far as order of indices.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3042665,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-11-11T16:54:09.907000",
              "content": "<p>baseline model: 0.001<br>\nflip z: lb=0<br>\ntranspose x,y: lb=0<br>\nremove 10 scale: lb=0<br>\n2 class (background versus particle): 0.003<br>\nraise error if average num of detected particle is less than 10:  in submission</p>\n<p>i think i know where we are getting to …</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 3041170,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2024-11-10T02:19:52.273000",
          "content": "<p>\" time alone exceeded 12 hours, \"<br>\nyou can just run DeepFindET e.g. for 50% of the test sample (then we expect a lower lb score, but not zero).</p>\n<p>But I think there is issue with the evaluation server and your local cv0.60 is correct</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3042196,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-11-11T09:56:42.253000",
      "content": "<p>the context of conversion is important.  (omitted because it is too long).<br>\nbasically I ask him lots of question of SNR, etc … those question that I do not understand from reading papers before asking him the kaggle question.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F47093115e66cb51f4f31af7b2d65e2f8%2FSelection_665.png?generation=1731318874920460&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F9f770d39b5398a85e582c52105306ceb%2FSelection_666.png?generation=1731319233851404&amp;alt=media\" alt=\"\"></p>",
      "votes": 1,
      "replies": [
        {
          "id": 3043065,
          "author_name": "Oleg Khudyakov",
          "author_url": "",
          "post_date": "2024-11-12T04:23:43.757000",
          "content": "<p>what is \"SNR\"? explain, please.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3043136,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-11-12T05:42:56.043000",
              "content": "<p>signal to noise ratio</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3041855,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-11-11T00:11:43.093000",
      "content": "<p>does the public test data contins simulated data only?<br>\n(and private contain both simulated and real data)</p>\n<p>we are given 7 real train data (and ask to use the  simulated data from the portal ourselves)<br>\nis this the explanation for zero score????</p>\n<hr>\n<p>please note the information from the dataset paper:</p>\n<p>Quotes: </p>\n<ol>\n<li><p>Dataset split: We curated a selection of <strong>492</strong> good-quality tomograms (Extended Data Fig. 7)<br>\nfrom the phantom dataset and divided them into three subsets: training, testing, and validation.</p></li>\n<li><p>All training data described here (7 tomograms) as well as additional published training data are<br>\navailable on the CZ CryoET Data Portal (CZCDP16, cryoetdataportal.czscience.com) under<br>\ndataset IDs 10440 (experimental data) and 10441 (PolNet43-simulations).</p></li>\n</ol>\n<hr>\n<p><a href=\"https://cryoetdataportal.czscience.com/browse-data/datasets\" target=\"_blank\">https://cryoetdataportal.czscience.com/browse-data/datasets</a><br>\nDataset ID: DS-10440 (7 runs)</p>\n<ul>\n<li>The data was acquired on a Krios G4 using a Falcon 4i detector and SelctrisX energy filter. Tomograms were reconstructed using AreTomo3 v1.0.23 and post-processed using different methods. Each run provides raw, denoised, missing wedge corrected and ctf corrected tomograms. </li>\n</ul>\n<p>Dataset ID: DS-10441 (27 runs)</p>\n<ul>\n<li>simulated tiltseries, tomograms and ground truth annotations for the purpose of training object identification algorithms in the context of the CryoET Object Identification Challenge. The data was simulated using polnet, and contains point labels</li>\n</ul>\n<hr>\n<p>now in  <a href=\"https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks/blob/main/DeepFindET/train.ipynb:\" target=\"_blank\">https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks/blob/main/DeepFindET/train.ipynb:</a></p>\n<p>Quotes: </p>\n<ul>\n<li>we demonstrate how to utilize this infrastructure to predict the 3D coordinates of six proteins of varying sizes, provided by the CryoET Dataportal (Dataset ID: 10439).  …</li>\n<li>copick config</li>\n</ul>\n<pre><code> : ,\n : ,\n   \n: ,\n            : ,\n            : 8,\n</code></pre>\n<p>is  CryoET Dataportal (Dataset ID: 10439) a miske? shouldn't it be 10441????</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3042025,
          "author_name": "coderRKJ",
          "author_url": "",
          "post_date": "2024-11-11T05:02:30.577000",
          "content": "<p>Looking at the <strong>Annotated Objects</strong> given in <a href=\"https://cryoetdataportal.czscience.com/datasets/10439\" target=\"_blank\">10439</a> and <a href=\"https://cryoetdataportal.czscience.com/datasets/10441\" target=\"_blank\">10441</a>, you seem to be right, it should be 10441.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3043622,
              "author_name": "Kyle Harrington",
              "author_url": "",
              "post_date": "2024-11-12T14:41:25.003000",
              "content": "<p>Ah, sorry about the dataset ID mix up there. When making the notebooks we used a different dataset to ensure that no data/info about the competition leaked, so we had an alternative synthetic dataset. You are correct 10441 is the synthetic dataset to use.</p>",
              "votes": 3,
              "replies": []
            }
          ]
        },
        {
          "id": 3043626,
          "author_name": "Kyle Harrington",
          "author_url": "",
          "post_date": "2024-11-12T14:48:35.063000",
          "content": "<p>to clarify this:</p>\n<pre><code>does the  test  contins simulated  ?\n(and  contain both simulated and  )\n\nwe are given   train  (and ask to  the simulated  from the portal ourselves)\nis this the explanation for zero score????\n</code></pre>\n<p>Public and private test data <em>only</em> contain real data. Simulated data is only provided to supplement training data.</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 3041171,
      "author_name": "David List",
      "author_url": "",
      "post_date": "2024-11-10T02:22:45.910000",
      "content": "<p>Interesting.  I submitted a ribosome only model that should have found at least something, but didn't even move up the leaderboard with a better 0.000 than before. 🤔</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3041179,
          "author_name": "David List",
          "author_url": "",
          "post_date": "2024-11-10T02:41:59.327000",
          "content": "<p>Oh, wait!  I <strong>did</strong> get a better 0.000.  Nice!  I'm still in the money!  😀</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10704200%2F467cb35e09122bd7226b49348b64e64c%2Fkaggle.jpg?generation=1731206452204999&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3041082,
      "author_name": "Andrei Zamfir",
      "author_url": "",
      "post_date": "2024-11-09T22:48:35.997000",
      "content": "<p>My disparity is far from yours, but I also experienced it nonetheless: local CV of ~0.013 -&gt; LB of ~0.000.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3041364,
      "author_name": "DennisSakva",
      "author_url": "",
      "post_date": "2024-11-10T08:20:04.060000",
      "content": "<p>Hi, Heng</p>\n<p>Your scaling \"10*coord and submit\" is a bit off.<br>\nCorrect scales are \"scale\": [10.012444196428572, 10.012444196428572, 10.012444537618887]<br>\nThe difference is tiny though</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3041392,
          "author_name": "David List",
          "author_url": "",
          "post_date": "2024-11-10T09:16:24.163000",
          "content": "<p><a href=\"https://www.kaggle.com/sakvaua\" target=\"_blank\">@sakvaua</a> Where are you getting that 10.012444196428572 from?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3041415,
              "author_name": "Andrei Zamfir",
              "author_url": "",
              "post_date": "2024-11-10T10:14:17.840000",
              "content": "<p>Host confirmed in <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/544895#3040071\" target=\"_blank\">here</a>, based on zarrs.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3041416,
              "author_name": "DennisSakva",
              "author_url": "",
              "post_date": "2024-11-10T10:15:01.363000",
              "content": "<p>Zarr attributes show this:<br>\n{\"multiscales\": [{\"axes\": [{\"name\": \"z\", \"type\": \"space\", \"unit\": \"angstrom\"}, {\"name\": \"y\", \"type\": \"space\", \"unit\": \"angstrom\"}, {\"name\": \"x\", \"type\": \"space\", \"unit\": \"angstrom\"}], \"datasets\": [{\"coordinateTransformations\": [{\"scale\": [10.012444196428572, 10.012444196428572, 10.012444537618887], \"type\": \"scale\"}], \"path\": \"0\"}, {\"coordinateTransformations\": [{\"scale\": [20.024888392857143, 20.024888392857143, 20.024889075237773], \"type\": \"scale\"}], \"path\": \"1\"}, {\"coordinateTransformations\": [{\"scale\": [40.049776785714286, 40.049776785714286, 40.04977815047555], \"type\": \"scale\"}], \"path\": \"2\"}], \"metadata\": {}, \"name\": \"/\", \"version\": \"0.4\"}]}</p>\n<p>scales for multiscale zars are 10.012, 20.024 and 40.049<br>\nI hope I'm not misinterpreting this</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 3041448,
              "author_name": "David List",
              "author_url": "",
              "post_date": "2024-11-10T11:21:17.323000",
              "content": "<p>Thanks!  That's pretty compelling.  😀</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3047078,
              "author_name": "David List",
              "author_url": "",
              "post_date": "2024-11-16T08:15:35.667000",
              "content": "<p>Actually, the scaling does start to matter for the apo-ferritin which only has a half-radius of 3 pixels.  You can visually see them off center if you don't do the scaling right.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 3041845,
          "author_name": "David List",
          "author_url": "",
          "post_date": "2024-11-10T23:53:07.613000",
          "content": "<p>7.56 units worst case (0.012 x 630).  Tiny, but might be the difference for some of the smaller particles in the lower right corner.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 3047079,
          "author_name": "David List",
          "author_url": "",
          "post_date": "2024-11-16T08:17:33.393000",
          "content": "<p>Looks like it actually starts to matter for the apo-ferratin with a half radius of only 3 pixels.  You can visually see them off center without the scaling.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3078340,
      "author_name": "hukaixin",
      "author_url": "",
      "post_date": "2024-12-22T07:23:08.717000",
      "content": "<p>Thank you for sharing, it has been very helpful to me</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3044900,
      "author_name": "David List",
      "author_url": "",
      "post_date": "2024-11-14T00:20:28.630000",
      "content": "<p>Pure speculation, but I keep getting exceptions/scoring errors when I submit despite the fact that everything works perfectly fine on the provided test/training data.  Wonder if maybe everyone is seeing that same issue but for most it causes only a few files to be processed instead of the exception I'm seeing.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3044901,
          "author_name": "David List",
          "author_url": "",
          "post_date": "2024-11-14T00:26:57.357000",
          "content": "<p>Scratch that.  I was iterating the test files but loading from the training directory.  😐</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3044703,
      "author_name": "chemdatafarmer",
      "author_url": "",
      "post_date": "2024-11-13T18:06:13.700000",
      "content": "<p>It would be nice if <a href=\"https://www.kaggle.com/kharrington\" target=\"_blank\">@kharrington</a> and/or Kaggle staff could confirm here that the organizers recognize something weird is going on and they are looking into it.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3043765,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-11-12T16:57:47.167000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3041217,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-11-10T03:22:47.047000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3041016": "I trained on 6 volumes and used the remaining one as validation. I could reach around local cv of about 0.35 to 0.75. However, the submitted lb score is always near 0.\n\nI suspect there is something wrong?\n\nmy approach:\n1. use only the highest resolution volume array:  [184, 630, 630]\n2. locate particle object in the array e.g. coord =(z,y,x)\n3. convert to location = 10*coord and submit\n\nevaluation script:\nhttps://www.kaggle.com/code/metric/czi-cryoet-84969\n\ndid you guys get the same funny results?",
    "3044890": "Howdy all ( @hengck23 @davidlist @cdeotte @sacuscreed @hypocrites @chemdatafarmer @herminw )\n\nWe have identified the bug in evaluation and will update soon.\n\nThank you for all of your investigations!\nKyle",
    "3043866": "Why is LB benchmark 0.587 and top Kaggler only have 0.003 after 1 week? This seems like a bug in Kaggle's scoring system. Has anyone contacted Kaggle about this?",
    "3041050": "It does look like something is weird since the leaderboard has a benchmark at 0.587 but the best submission is only 0.002",
    "3044346": "this is something interesting: local cv score for randomly generated 100 points:\n\n```\n\n['TS_5_4', 'TS_6_4', 'TS_6_6', 'TS_69_2', 'TS_73_6', 'TS_86_3', 'TS_99_9']\nlb_score 0.0026500562150282825\n         particle_type    P    T  hit  miss   fp  precision    recall   f-beta4  weight\n0         apo-ferritin  700  375    0   375  700   0.000000  0.000000  0.000000       1\n1         beta-amylase  700   87    0    87  700   0.000000  0.000000  0.000000       0\n2   beta-galactosidase  700  112    0   112  700   0.000000  0.000000  0.000000       2\n3             ribosome  700  331    4   327  696   0.005714  0.012085  0.011341       1\n4        thyroglobulin  700  251    1   250  699   0.001429  0.003984  0.003605       2\n5  virus-like-particle  700  113    0   113  700   0.000000  0.000000  0.000000       1\n\n```\n\n\n```\nsubmit_df=[]\nfor id in SIMPLE_ID:\n    for n in PARTICLE_NAME:\n        print(id,n)\n        D,H,W = 184, 630, 630 \n        xyz=np.random.uniform(0,1,(100,3)) \n        xyz = xyz*[[W,H,0.5*D]]+[[0,0,0.25*D]]\n        xyz = xyz*10 \n        submit_df.append(\n            pd.DataFrame({'experiment': id, 'particle_type': n, 'x': xyz[:, 0], 'y': xyz[:, 1], 'z': xyz[:, 2]})\n        ) \nsubmit_df = pd.concat(submit_df)\nsubmit_df.insert(loc=0, column='id', value=np.arange(len(submit_df)))\nprint(submit_df)\n\n```",
    "3042243": "@kharrington what do you think about this? Seems there is something weird with submission evaluation.",
    "3041167": "The sample submission notebook using blob detector: indeed zero for local lb\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fbaf2c5967cb015e9ad1417dd38f93857%2FSelection_657.png?generation=1731204927157155&alt=media)",
    "3041126": "Similarly, cv0.6 lb0.002, and I believe that the benchmark on the leaderboard was not evaluated through Kaggle's notebook, it is possible that only the CSV was uploaded. I have also tried DeepFindET for benchmarks, and the inference time alone exceeded 12 hours,  The time of coordinate conversion is not included",
    "3042196": "the context of conversion is important.  (omitted because it is too long).\nbasically I ask him lots of question of SNR, etc ... those question that I do not understand from reading papers before asking him the kaggle question.\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F47093115e66cb51f4f31af7b2d65e2f8%2FSelection_665.png?generation=1731318874920460&alt=media)\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F9f770d39b5398a85e582c52105306ceb%2FSelection_666.png?generation=1731319233851404&alt=media)",
    "3041855": "does the public test data contins simulated data only?\n(and private contain both simulated and real data)\n\nwe are given 7 real train data (and ask to use the  simulated data from the portal ourselves)\nis this the explanation for zero score????\n\n---\n\nplease note the information from the dataset paper:\n\nQuotes: \n\n1. Dataset split: We curated a selection of **492** good-quality tomograms (Extended Data Fig. 7)\nfrom the phantom dataset and divided them into three subsets: training, testing, and validation.\n\n2. All training data described here (7 tomograms) as well as additional published training data are\navailable on the CZ CryoET Data Portal (CZCDP16, cryoetdataportal.czscience.com) under\ndataset IDs 10440 (experimental data) and 10441 (PolNet43-simulations).\n\n---\n\nhttps://cryoetdataportal.czscience.com/browse-data/datasets\nDataset ID: DS-10440 (7 runs)\n- The data was acquired on a Krios G4 using a Falcon 4i detector and SelctrisX energy filter. Tomograms were reconstructed using AreTomo3 v1.0.23 and post-processed using different methods. Each run provides raw, denoised, missing wedge corrected and ctf corrected tomograms. \n\n\nDataset ID: DS-10441 (27 runs)\n- simulated tiltseries, tomograms and ground truth annotations for the purpose of training object identification algorithms in the context of the CryoET Object Identification Challenge. The data was simulated using polnet, and contains point labels\n\n\n---\n\nnow in  https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks/blob/main/DeepFindET/train.ipynb:\n\n\nQuotes: \n\n-  we demonstrate how to utilize this infrastructure to predict the 3D coordinates of six proteins of varying sizes, provided by the CryoET Dataportal (Dataset ID: 10439).  ...\n-  copick config\n```\n \"name\": \"apo-ferritin\",\n \"name\": \"beta-amylase\",\n...   \n\"name\": \"membrane\",\n            \"is_particle\": false,\n            \"label\": 8,\n\n```\n\nis  CryoET Dataportal (Dataset ID: 10439) a miske? shouldn't it be 10441????\n",
    "3041171": "Interesting.  I submitted a ribosome only model that should have found at least something, but didn't even move up the leaderboard with a better 0.000 than before. 🤔",
    "3041082": "My disparity is far from yours, but I also experienced it nonetheless: local CV of ~0.013 -> LB of ~0.000.",
    "3041364": "Hi, Heng\n\nYour scaling \"10*coord and submit\" is a bit off.\nCorrect scales are \"scale\": [10.012444196428572, 10.012444196428572, 10.012444537618887]\nThe difference is tiny though",
    "3078340": "Thank you for sharing, it has been very helpful to me",
    "3044900": "Pure speculation, but I keep getting exceptions/scoring errors when I submit despite the fact that everything works perfectly fine on the provided test/training data.  Wonder if maybe everyone is seeing that same issue but for most it causes only a few files to be processed instead of the exception I'm seeing.",
    "3044703": "It would be nice if @kharrington and/or Kaggle staff could confirm here that the organizers recognize something weird is going on and they are looking into it.",
    "3043765": "",
    "3041217": ""
  }
}