{
  "id": 545734,
  "title": "Very new to everything - Just exploring - But I wondered about this....",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/545734",
  "author_name": "",
  "post_date": "2024-11-11T21:44:33.596457600Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I am very new to this, but I tinker with generated images and code as I am trying to learn.</p>\n<p>Piecing together I wanted to know, about any value in rather than trying to detect was was desired, what about ruling out what wasn't, thus only leaving particles of potential interest.</p>\n<p>I have used the TS_5_4/VoxelSpacing10.000/ctfdeconvolved.zarr with all picked data…. Setting the Z value to 1-9999 to capture all possible particles of interest across all 184 images, So even when they are not \"visible\", at least their disturbance might be?</p>\n<p>I have cut the regions to a 20by20 pixel region, which would be contained in the circle, and used the clean image on the left to make the cut to avoid the bounding region.</p>\n<p>I have followed every feature across all the images.</p>\n<p>I then generated 300 random 20 by 20 pixel regions, which were at least 3 pixels from a region of interest, and followed these for all 184 images, which I named nOi.</p>\n<p>I then applied Kmeans clustering, to see how the images would cluster, base don there being an underlying feature that may be of interest, which wouldn't be detected, but the \"disturbance\" might be…</p>\n<p>here are my results, the image containing the particles of interest, and the 300 random regions which were cut out.</p>\n<p>I wondered if this approach had any merit, or explain why this doesn't work.</p>\n<p>[The images are of, The data table of kmeans results, the coloured circle image shows all particles of interest and their location, the black squared image shows where my 300 random regions were cut from]</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18343280%2Fe110ca8eb50210b3992b450c634c1090%2FTS_5_4_all_001%20-%20Copy.png?generation=1731360841146622&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18343280%2F217bcc1dc8c74ff4662caa4cb7a52a7f%2FNoI%20-%20Copy.png?generation=1731360858090464&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18343280%2F1fac15cb958c62d9e02093563935aa2a%2Fdata_table%20-%20Copy.png?generation=1731360872811751&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "3042897",
      "postDate": "11/11/2024 21:44:33",
      "content": "<p>I am very new to this, but I tinker with generated images and code as I am trying to learn.</p>\n<p>Piecing together I wanted to know, about any value in rather than trying to detect was was desired, what about ruling out what wasn't, thus only leaving particles of potential interest.</p>\n<p>I have used the TS_5_4/VoxelSpacing10.000/ctfdeconvolved.zarr with all picked data…. Setting the Z value to 1-9999 to capture all possible particles of interest across all 184 images, So even when they are not \"visible\", at least their disturbance might be?</p>\n<p>I have cut the regions to a 20by20 pixel region, which would be contained in the circle, and used the clean image on the left to make the cut to avoid the bounding region.</p>\n<p>I have followed every feature across all the images.</p>\n<p>I then generated 300 random 20 by 20 pixel regions, which were at least 3 pixels from a region of interest, and followed these for all 184 images, which I named nOi.</p>\n<p>I then applied Kmeans clustering, to see how the images would cluster, base don there being an underlying feature that may be of interest, which wouldn't be detected, but the \"disturbance\" might be…</p>\n<p>here are my results, the image containing the particles of interest, and the 300 random regions which were cut out.</p>\n<p>I wondered if this approach had any merit, or explain why this doesn't work.</p>\n<p>[The images are of, The data table of kmeans results, the coloured circle image shows all particles of interest and their location, the black squared image shows where my 300 random regions were cut from]</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18343280%2Fe110ca8eb50210b3992b450c634c1090%2FTS_5_4_all_001%20-%20Copy.png?generation=1731360841146622&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18343280%2F217bcc1dc8c74ff4662caa4cb7a52a7f%2FNoI%20-%20Copy.png?generation=1731360858090464&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18343280%2F1fac15cb958c62d9e02093563935aa2a%2Fdata_table%20-%20Copy.png?generation=1731360872811751&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I am very new to this, but I tinker with generated images and code as I am trying to learn.\n\nPiecing together I wanted to know, about any value in rather than trying to detect was was desired, what about ruling out what wasn't, thus only leaving particles of potential interest.\n\nI have used the TS_5_4/VoxelSpacing10.000/ctfdeconvolved.zarr with all picked data.... Setting the Z value to 1-9999 to capture all possible particles of interest across all 184 images, So even when they are not \"visible\", at least their disturbance might be?\n\nI have cut the regions to a 20by20 pixel region, which would be contained in the circle, and used the clean image on the left to make the cut to avoid the bounding region.\n\nI have followed every feature across all the images.\n\nI then generated 300 random 20 by 20 pixel regions, which were at least 3 pixels from a region of interest, and followed these for all 184 images, which I named nOi.\n\nI then applied Kmeans clustering, to see how the images would cluster, base don there being an underlying feature that may be of interest, which wouldn't be detected, but the \"disturbance\" might be...\n\nhere are my results, the image containing the particles of interest, and the 300 random regions which were cut out.\n\nI wondered if this approach had any merit, or explain why this doesn't work.\n\n[The images are of, The data table of kmeans results, the coloured circle image shows all particles of interest and their location, the black squared image shows where my 300 random regions were cut from]\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18343280%2Fe110ca8eb50210b3992b450c634c1090%2FTS_5_4_all_001%20-%20Copy.png?generation=1731360841146622&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18343280%2F217bcc1dc8c74ff4662caa4cb7a52a7f%2FNoI%20-%20Copy.png?generation=1731360858090464&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18343280%2F1fac15cb958c62d9e02093563935aa2a%2Fdata_table%20-%20Copy.png?generation=1731360872811751&alt=media)",
      "votes": null
    },
    {
      "id": "3043999",
      "postDate": "11/12/2024 22:30:49",
      "content": "<p>Hi Boris, your approach looks interesting. But I don't want to comment on your approach, rather the results you have shown. The tomogram slice you are showing seems to be at the top/bottom edge of the sample. I can understand your results better if you show a Z=~90 slice that is in the middle of the tomogram. Thank you.</p>",
      "rawMarkdown": "Hi Boris, your approach looks interesting. But I don't want to comment on your approach, rather the results you have shown. The tomogram slice you are showing seems to be at the top/bottom edge of the sample. I can understand your results better if you show a Z=~90 slice that is in the middle of the tomogram. Thank you.",
      "votes": null
    },
    {
      "id": "3044029",
      "postDate": "11/12/2024 23:15:55",
      "content": "<p>you can google or ask chatgpt on learning a model (which should be an embedding one that learns distance) to do k-means clustering.<br>\nthe loss should be the difference of the model-produced clusters and the ground truth clusters.</p>\n<p>since you are likey to produce many candidates, a ranking output (indicating if the candidate is background or not) is useful.</p>\n<hr>\n<p>now if we make use of some tricks:</p>\n<ol>\n<li>we assume (is this correct?) that every (or almost all) test volume will have at least some particles of the 6 class/cluster. then our task is to select/or rank each voxel  if they belong to the class or not</li>\n<li>this is softmax over all voxel z,y,x. for each class</li>\n</ol>\n<p>(it is not the same as the softmax over class for each volume, the usual segmentation of voxel)</p>",
      "rawMarkdown": "you can google or ask chatgpt on learning a model (which should be an embedding one that learns distance) to do k-means clustering.\nthe loss should be the difference of the model-produced clusters and the ground truth clusters.\n\nsince you are likey to produce many candidates, a ranking output (indicating if the candidate is background or not) is useful.\n\n---\n\nnow if we make use of some tricks:\n\n1. we assume (is this correct?) that every (or almost all) test volume will have at least some particles of the 6 class/cluster. then our task is to select/or rank each voxel  if they belong to the class or not\n2. this is softmax over all voxel z,y,x. for each class\n\n(it is not the same as the softmax over class for each volume, the usual segmentation of voxel)",
      "votes": null
    },
    {
      "id": "3044593",
      "postDate": "11/13/2024 15:15:30",
      "content": "<p>Hi,  I am not entirely sure how to visualise that at this moment, I copied the basis of the code from David List, which was his ribosome visualization, I just overlayed all particles on it, and set the Z to 1-9999.</p>\n<p>Every slice that was made I cut out the position of the known molecule, whether it was actually visible or not. So my random chunks, may actually be just nothing all the way through, or they may hit different cellular contents.</p>\n<p>Likewise, the particles of interest, are what the region looks like knowing there is a particle of interest in that space, but it isn't currently detected.</p>\n<p>But continuing on, from what I posted above…. Using the K-means clustering with 25 groups.  <br>\nThe groups the images fall into would suggest a method of unique identification, based on their \"positive\" or \"negative\" selection?</p>\n<p>This is however still just TS_5_4, but it looks promising?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18343280%2F0747751ed6a1acf2da89f6e9ce47ce6c%2F25_clusters%20-%20Copy.png?generation=1731510164147235&amp;alt=media\" alt=\"\"></p>\n<p>The first data table is 2 STdev Above average in green, and 2 STdev below in red<br>\nThe 2nd is with an arbitrary   20% above or -20% below <br>\nThe 3rd is Relevant groups only, <br>\nThe 4th summarizes the groups,  suggesting the unique footprint, of the groups the images belonged to.<br>\nFor Example, VirusLikeParticles, would clearly be Group 0, and Apoferritin, clearly group 4.<br>\nbeta-amylase suggested as impossible, has a very clear positive of 14 and 19, with large negs at 1 and 23. <br>\nBut the combinations, may make for some method of identification.</p>\n<p>For context</p>\n<p>Slice 50:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18343280%2Fc4038bbef89d1e5de163d2ee78d73f65%2FTS_5_4_all_050%20-%20Copy.png?generation=1731510802317763&amp;alt=media\" alt=\"\"></p>\n<p>slice 90:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18343280%2Fda90ddb43ebb0e3e00487f436781044c%2FTS_5_4_all_090%20-%20Copy.png?generation=1731510836662429&amp;alt=media\" alt=\"\"></p>\n<p>slice 130:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18343280%2F9e75d1843ada211d50b242469b1534b3%2FTS_5_4_all_130%20-%20Copy.png?generation=1731510861164270&amp;alt=media\" alt=\"\"></p>\n<p>The random segments were also constant all the way through</p>",
      "rawMarkdown": "Hi,  I am not entirely sure how to visualise that at this moment, I copied the basis of the code from David List, which was his ribosome visualization, I just overlayed all particles on it, and set the Z to 1-9999.\n\nEvery slice that was made I cut out the position of the known molecule, whether it was actually visible or not. So my random chunks, may actually be just nothing all the way through, or they may hit different cellular contents.\n\nLikewise, the particles of interest, are what the region looks like knowing there is a particle of interest in that space, but it isn't currently detected.\n\nBut continuing on, from what I posted above.... Using the K-means clustering with 25 groups.  \nThe groups the images fall into would suggest a method of unique identification, based on their \"positive\" or \"negative\" selection?\n\nThis is however still just TS_5_4, but it looks promising?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18343280%2F0747751ed6a1acf2da89f6e9ce47ce6c%2F25_clusters%20-%20Copy.png?generation=1731510164147235&alt=media)\n\nThe first data table is 2 STdev Above average in green, and 2 STdev below in red\nThe 2nd is with an arbitrary   20% above or -20% below \nThe 3rd is Relevant groups only, \nThe 4th summarizes the groups,  suggesting the unique footprint, of the groups the images belonged to.\nFor Example, VirusLikeParticles, would clearly be Group 0, and Apoferritin, clearly group 4.\nbeta-amylase suggested as impossible, has a very clear positive of 14 and 19, with large negs at 1 and 23. \nBut the combinations, may make for some method of identification.\n\nFor context\n\nSlice 50:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18343280%2Fc4038bbef89d1e5de163d2ee78d73f65%2FTS_5_4_all_050%20-%20Copy.png?generation=1731510802317763&alt=media)\n\nslice 90:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18343280%2Fda90ddb43ebb0e3e00487f436781044c%2FTS_5_4_all_090%20-%20Copy.png?generation=1731510836662429&alt=media)\n\nslice 130:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18343280%2F9e75d1843ada211d50b242469b1534b3%2FTS_5_4_all_130%20-%20Copy.png?generation=1731510861164270&alt=media)\n\nThe random segments were also constant all the way through",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3043999,
      "author_name": "rezaparaan",
      "author_url": "",
      "post_date": "11/12/2024 22:30:49",
      "content": "<p>Hi Boris, your approach looks interesting. But I don't want to comment on your approach, rather the results you have shown. The tomogram slice you are showing seems to be at the top/bottom edge of the sample. I can understand your results better if you show a Z=~90 slice that is in the middle of the tomogram. Thank you.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3044593,
          "author_name": "borismclane",
          "author_url": "",
          "post_date": "11/13/2024 15:15:30",
          "content": "<p>Hi,  I am not entirely sure how to visualise that at this moment, I copied the basis of the code from David List, which was his ribosome visualization, I just overlayed all particles on it, and set the Z to 1-9999.</p>\n<p>Every slice that was made I cut out the position of the known molecule, whether it was actually visible or not. So my random chunks, may actually be just nothing all the way through, or they may hit different cellular contents.</p>\n<p>Likewise, the particles of interest, are what the region looks like knowing there is a particle of interest in that space, but it isn't currently detected.</p>\n<p>But continuing on, from what I posted above…. Using the K-means clustering with 25 groups.  <br>\nThe groups the images fall into would suggest a method of unique identification, based on their \"positive\" or \"negative\" selection?</p>\n<p>This is however still just TS_5_4, but it looks promising?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18343280%2F0747751ed6a1acf2da89f6e9ce47ce6c%2F25_clusters%20-%20Copy.png?generation=1731510164147235&amp;alt=media\" alt=\"\"></p>\n<p>The first data table is 2 STdev Above average in green, and 2 STdev below in red<br>\nThe 2nd is with an arbitrary   20% above or -20% below <br>\nThe 3rd is Relevant groups only, <br>\nThe 4th summarizes the groups,  suggesting the unique footprint, of the groups the images belonged to.<br>\nFor Example, VirusLikeParticles, would clearly be Group 0, and Apoferritin, clearly group 4.<br>\nbeta-amylase suggested as impossible, has a very clear positive of 14 and 19, with large negs at 1 and 23. <br>\nBut the combinations, may make for some method of identification.</p>\n<p>For context</p>\n<p>Slice 50:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18343280%2Fc4038bbef89d1e5de163d2ee78d73f65%2FTS_5_4_all_050%20-%20Copy.png?generation=1731510802317763&amp;alt=media\" alt=\"\"></p>\n<p>slice 90:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18343280%2Fda90ddb43ebb0e3e00487f436781044c%2FTS_5_4_all_090%20-%20Copy.png?generation=1731510836662429&amp;alt=media\" alt=\"\"></p>\n<p>slice 130:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18343280%2F9e75d1843ada211d50b242469b1534b3%2FTS_5_4_all_130%20-%20Copy.png?generation=1731510861164270&amp;alt=media\" alt=\"\"></p>\n<p>The random segments were also constant all the way through</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3044029,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "11/12/2024 23:15:55",
      "content": "<p>you can google or ask chatgpt on learning a model (which should be an embedding one that learns distance) to do k-means clustering.<br>\nthe loss should be the difference of the model-produced clusters and the ground truth clusters.</p>\n<p>since you are likey to produce many candidates, a ranking output (indicating if the candidate is background or not) is useful.</p>\n<hr>\n<p>now if we make use of some tricks:</p>\n<ol>\n<li>we assume (is this correct?) that every (or almost all) test volume will have at least some particles of the 6 class/cluster. then our task is to select/or rank each voxel  if they belong to the class or not</li>\n<li>this is softmax over all voxel z,y,x. for each class</li>\n</ol>\n<p>(it is not the same as the softmax over class for each volume, the usual segmentation of voxel)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3042897": "I am very new to this, but I tinker with generated images and code as I am trying to learn.\n\nPiecing together I wanted to know, about any value in rather than trying to detect was was desired, what about ruling out what wasn't, thus only leaving particles of potential interest.\n\nI have used the TS_5_4/VoxelSpacing10.000/ctfdeconvolved.zarr with all picked data.... Setting the Z value to 1-9999 to capture all possible particles of interest across all 184 images, So even when they are not \"visible\", at least their disturbance might be?\n\nI have cut the regions to a 20by20 pixel region, which would be contained in the circle, and used the clean image on the left to make the cut to avoid the bounding region.\n\nI have followed every feature across all the images.\n\nI then generated 300 random 20 by 20 pixel regions, which were at least 3 pixels from a region of interest, and followed these for all 184 images, which I named nOi.\n\nI then applied Kmeans clustering, to see how the images would cluster, base don there being an underlying feature that may be of interest, which wouldn't be detected, but the \"disturbance\" might be...\n\nhere are my results, the image containing the particles of interest, and the 300 random regions which were cut out.\n\nI wondered if this approach had any merit, or explain why this doesn't work.\n\n[The images are of, The data table of kmeans results, the coloured circle image shows all particles of interest and their location, the black squared image shows where my 300 random regions were cut from]\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18343280%2Fe110ca8eb50210b3992b450c634c1090%2FTS_5_4_all_001%20-%20Copy.png?generation=1731360841146622&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18343280%2F217bcc1dc8c74ff4662caa4cb7a52a7f%2FNoI%20-%20Copy.png?generation=1731360858090464&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18343280%2F1fac15cb958c62d9e02093563935aa2a%2Fdata_table%20-%20Copy.png?generation=1731360872811751&alt=media)",
    "3043999": "Hi Boris, your approach looks interesting. But I don't want to comment on your approach, rather the results you have shown. The tomogram slice you are showing seems to be at the top/bottom edge of the sample. I can understand your results better if you show a Z=~90 slice that is in the middle of the tomogram. Thank you.",
    "3044029": "you can google or ask chatgpt on learning a model (which should be an embedding one that learns distance) to do k-means clustering.\nthe loss should be the difference of the model-produced clusters and the ground truth clusters.\n\nsince you are likey to produce many candidates, a ranking output (indicating if the candidate is background or not) is useful.\n\n---\n\nnow if we make use of some tricks:\n\n1. we assume (is this correct?) that every (or almost all) test volume will have at least some particles of the 6 class/cluster. then our task is to select/or rank each voxel  if they belong to the class or not\n2. this is softmax over all voxel z,y,x. for each class\n\n(it is not the same as the softmax over class for each volume, the usual segmentation of voxel)",
    "3044593": "Hi,  I am not entirely sure how to visualise that at this moment, I copied the basis of the code from David List, which was his ribosome visualization, I just overlayed all particles on it, and set the Z to 1-9999.\n\nEvery slice that was made I cut out the position of the known molecule, whether it was actually visible or not. So my random chunks, may actually be just nothing all the way through, or they may hit different cellular contents.\n\nLikewise, the particles of interest, are what the region looks like knowing there is a particle of interest in that space, but it isn't currently detected.\n\nBut continuing on, from what I posted above.... Using the K-means clustering with 25 groups.  \nThe groups the images fall into would suggest a method of unique identification, based on their \"positive\" or \"negative\" selection?\n\nThis is however still just TS_5_4, but it looks promising?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18343280%2F0747751ed6a1acf2da89f6e9ce47ce6c%2F25_clusters%20-%20Copy.png?generation=1731510164147235&alt=media)\n\nThe first data table is 2 STdev Above average in green, and 2 STdev below in red\nThe 2nd is with an arbitrary   20% above or -20% below \nThe 3rd is Relevant groups only, \nThe 4th summarizes the groups,  suggesting the unique footprint, of the groups the images belonged to.\nFor Example, VirusLikeParticles, would clearly be Group 0, and Apoferritin, clearly group 4.\nbeta-amylase suggested as impossible, has a very clear positive of 14 and 19, with large negs at 1 and 23. \nBut the combinations, may make for some method of identification.\n\nFor context\n\nSlice 50:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18343280%2Fc4038bbef89d1e5de163d2ee78d73f65%2FTS_5_4_all_050%20-%20Copy.png?generation=1731510802317763&alt=media)\n\nslice 90:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18343280%2Fda90ddb43ebb0e3e00487f436781044c%2FTS_5_4_all_090%20-%20Copy.png?generation=1731510836662429&alt=media)\n\nslice 130:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18343280%2F9e75d1843ada211d50b242469b1534b3%2FTS_5_4_all_130%20-%20Copy.png?generation=1731510861164270&alt=media)\n\nThe random segments were also constant all the way through"
  },
  "source": "meta"
}