{
  "id": 549531,
  "title": "What's your CV-LB combination?",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/549531",
  "author_name": "DennisSakva",
  "post_date": "2024-12-02T16:57:33.721000",
  "votes": 12,
  "comment_count": 31,
  "views": 0,
  "content": "<p>I'll start. Train - first five, validation on the last two<br>\nCV 0.8 LB 0.7 with poor correlation between CV and LB</p>",
  "messages": [
    {
      "id": 3061768,
      "postDate": "2024-12-03T01:14:09.827Z",
      "content": "<p>please imagine how the loss curves for train,validation and hidden test would look like over different numbers of train, validation and test samples.<br>\n(Hint: if you cannot imagine how the curve look like, download cifar10 and do experiments and plot the curve for different umbers of train,validation and  test. assumption here is that the data difficulty of cirfar10 is same as the kaggle competition, if not \"imagine\" the change in cifar curves to be accounted for when \"transfering them\" to this kaggle dataset)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc27daac31cf753118576333f50fe3ec1%2FSelection_198.png?generation=1733224281913535&amp;alt=media\" alt=\"\"></p>\n<p>there should be 4 curve (train,valid, hidden public/private test). i only show two (train, valid). you can probe for the rest.</p>\n<p>you are reporting only one black point on the curve. It is better to consider the whole curve itself (i.e. submitting over multiple models at different epochs)</p>\n<hr>\n<p>once you understand this concept, then \"shakeup or not\", \"lb/cv correlation\", etc is only statistics.</p>\n<p>here \"parameters\" are : num of train, num of valid samples, model complexity and we are trying to </p>\n<ol>\n<li>make hidden test curve overlaping with our valid curve as much as possible. </li>\n<li>select a point (i.e. how many epoch to train model) in the overlapping region for submission</li>\n</ol>",
      "rawMarkdown": "please imagine how the loss curves for train,validation and hidden test would look like over different numbers of train, validation and test samples.\n(Hint: if you cannot imagine how the curve look like, download cifar10 and do experiments and plot the curve for different umbers of train,validation and  test. assumption here is that the data difficulty of cirfar10 is same as the kaggle competition, if not \"imagine\" the change in cifar curves to be accounted for when \"transfering them\" to this kaggle dataset)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc27daac31cf753118576333f50fe3ec1%2FSelection_198.png?generation=1733224281913535&alt=media)\n\n\nthere should be 4 curve (train,valid, hidden public/private test). i only show two (train, valid). you can probe for the rest.\n\nyou are reporting only one black point on the curve. It is better to consider the whole curve itself (i.e. submitting over multiple models at different epochs)\n\n---\n\nonce you understand this concept, then \"shakeup or not\", \"lb/cv correlation\", etc is only statistics.\n\nhere \"parameters\" are : num of train, num of valid samples, model complexity and we are trying to \n1. make hidden test curve overlaping with our valid curve as much as possible. \n2. select a point (i.e. how many epoch to train model) in the overlapping region for submission",
      "votes": 14,
      "replies": [
        {
          "id": 3061957,
          "postDate": "2024-12-03T06:46:50.103Z",
          "content": "<p>Great thanks, your post resolves my confusion in this competition.</p>",
          "rawMarkdown": "Great thanks, your post resolves my confusion in this competition."
        },
        {
          "id": 3061966,
          "postDate": "2024-12-03T07:06:44.703Z",
          "content": "<p>Hi, Heng. Thanks for the explanation. I wasn’t looking for the theoretical reasons behind the difference. I just wanted to gather a few data points on what others are observing regarding the CV-LB discrepancy. From the few I’ve seen so far, the difference seems to consistently be around 0.1. Three data points aren’t enough for any definitive conclusions, but it’s at least something to think about! <br>\nCheers.</p>",
          "rawMarkdown": "Hi, Heng. Thanks for the explanation. I wasn’t looking for the theoretical reasons behind the difference. I just wanted to gather a few data points on what others are observing regarding the CV-LB discrepancy. From the few I’ve seen so far, the difference seems to consistently be around 0.1. Three data points aren’t enough for any definitive conclusions, but it’s at least something to think about! \nCheers.",
          "votes": 1,
          "replies": [
            {
              "id": 3061975,
              "postDate": "2024-12-03T07:19:04.377Z",
              "content": "<p>It could also be that we're all using the same tomogram for validation.  😀</p>",
              "rawMarkdown": "It could also be that we're all using the same tomogram for validation.  😀"
            },
            {
              "id": 3061983,
              "postDate": "2024-12-03T07:31:48.907Z",
              "content": "<p>Sure thing! But consistency is a good sign and allows to infer which files give better estimates of which particles.</p>",
              "rawMarkdown": "Sure thing! But consistency is a good sign and allows to infer which files give better estimates of which particles.",
              "votes": 1
            },
            {
              "id": 3061992,
              "postDate": "2024-12-03T07:39:12.490Z",
              "content": "<p>TS_6_4, which I think a lot of people are using may be the odd one of the 7, because it's the only one with the strange viruses.  I doubt that explains everything, but it might be a small piece…  Actually, scratch that.  If TS_6_4 is more like the LB data, that would have the opposite effect.<br>\n<img src=\"https://www.kaggleusercontent.com/kf/210412070/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..RNYcK5DUtlQkP555-N8opg.lnLUsGlsB9Ri3l5VCUUlLYTJ0NKlPmh90c2dO_ua-vvzPdjsvHfjDGBJT-be5BrHutGtL0ftq-o0xHLcqlfAQhr7Y38sjkrOh372vggex2nEG7Fq8O4UiRyaD-7v6_MnJHlkaSq_heDXXDGFsQjJFUbyafwDxIh_bSwJZLQrBzfNBDoBkJOBdpXKJi2hFifZxIlsoyXvLv8rheoUXC6lrQa8xl1klTL3DT71p_bYBGwR7l2_8e9ENFFEVOWA7uRAEdoc6Bo6A0KKA2ZjGCXzNcYNYjE4Vsa8c_65W48Ly7IUbeAboGd1B6f0LeKwzM4TN3NqG2jCOKRa2IsF_Y-6Idl9PhfaOP3_0LHp0SnU0Jzs4QTgD9Gi0TrnfNIXqjYuna9BsHR6KwQPZgc8B_8emlDLywb3UM3YsOXaLAUA8A9c637jc1tjZLMlVtPqWtHOn8zbpqIrNX7OynQZOGZcQeuBE1pIvcF51A8FT-PSCb4-bjknY_x2e142wHZCQiVhHIfFI0HDftNLoC6jjfFSZaMnFtZhaAeWWLGoVknSTDXNH189QObu-xwouffCB3QK84Sc87SCnrivKaaeA9U6ndshiwrjPCxdYU2RZKYsn4T4SiY7HHw4wXE-esPJEuLK551R36c4Z5bW-nsEaOm8jg.CyN4EmtJly4lEqLASAA8rA/__results___files/__results___9_0.png\" alt=\"\"></p>",
              "rawMarkdown": "TS_6_4, which I think a lot of people are using may be the odd one of the 7, because it's the only one with the strange viruses.  I doubt that explains everything, but it might be a small piece...  Actually, scratch that.  If TS_6_4 is more like the LB data, that would have the opposite effect.\n![](https://www.kaggleusercontent.com/kf/210412070/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..RNYcK5DUtlQkP555-N8opg.lnLUsGlsB9Ri3l5VCUUlLYTJ0NKlPmh90c2dO_ua-vvzPdjsvHfjDGBJT-be5BrHutGtL0ftq-o0xHLcqlfAQhr7Y38sjkrOh372vggex2nEG7Fq8O4UiRyaD-7v6_MnJHlkaSq_heDXXDGFsQjJFUbyafwDxIh_bSwJZLQrBzfNBDoBkJOBdpXKJi2hFifZxIlsoyXvLv8rheoUXC6lrQa8xl1klTL3DT71p_bYBGwR7l2_8e9ENFFEVOWA7uRAEdoc6Bo6A0KKA2ZjGCXzNcYNYjE4Vsa8c_65W48Ly7IUbeAboGd1B6f0LeKwzM4TN3NqG2jCOKRa2IsF_Y-6Idl9PhfaOP3_0LHp0SnU0Jzs4QTgD9Gi0TrnfNIXqjYuna9BsHR6KwQPZgc8B_8emlDLywb3UM3YsOXaLAUA8A9c637jc1tjZLMlVtPqWtHOn8zbpqIrNX7OynQZOGZcQeuBE1pIvcF51A8FT-PSCb4-bjknY_x2e142wHZCQiVhHIfFI0HDftNLoC6jjfFSZaMnFtZhaAeWWLGoVknSTDXNH189QObu-xwouffCB3QK84Sc87SCnrivKaaeA9U6ndshiwrjPCxdYU2RZKYsn4T4SiY7HHw4wXE-esPJEuLK551R36c4Z5bW-nsEaOm8jg.CyN4EmtJly4lEqLASAA8rA/__results___files/__results___9_0.png)",
              "votes": 2
            },
            {
              "id": 3062004,
              "postDate": "2024-12-03T07:52:37Z",
              "content": "<p>Sorry, that image is a bit longer than I remembered.  😀</p>",
              "rawMarkdown": "Sorry, that image is a bit longer than I remembered.  😀",
              "votes": 1
            },
            {
              "id": 3062106,
              "postDate": "2024-12-03T09:48:34.577Z",
              "content": "<p>But it's cool :)</p>",
              "rawMarkdown": "But it's cool :)",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3061477,
      "postDate": "2024-12-02T16:57:33.723Z",
      "content": "<p>I'll start. Train - first five, validation on the last two<br>\nCV 0.8 LB 0.7 with poor correlation between CV and LB</p>",
      "rawMarkdown": "I'll start. Train - first five, validation on the last two\nCV 0.8 LB 0.7 with poor correlation between CV and LB",
      "votes": 12
    },
    {
      "id": 3070925,
      "postDate": "2024-12-13T07:06:25.123Z",
      "content": "<p>The way Cross Validation is split affects the gap between CV and LB scores. However, up to a certain score threshold, it seems that CV and LB are correlated.<br>\nEssentially, I believe this phenomenon occurs when the training data includes samples similar to those used for calculating the Public Leaderboard, resulting in higher LB scores.</p>\n<p>CV0.77 LB0.7</p>",
      "rawMarkdown": "The way Cross Validation is split affects the gap between CV and LB scores. However, up to a certain score threshold, it seems that CV and LB are correlated.\nEssentially, I believe this phenomenon occurs when the training data includes samples similar to those used for calculating the Public Leaderboard, resulting in higher LB scores.\n\nCV0.77 LB0.7",
      "votes": 1
    },
    {
      "id": 3070012,
      "postDate": "2024-12-12T06:15:16.947Z",
      "content": "<p>Validation on TS_6_4<br>\nCV: 0.755 LB: 0.714<br>\nCV: 0.822 LB: 0.616<br>\nLooks like I overfit😂</p>",
      "rawMarkdown": "Validation on TS_6_4\nCV: 0.755 LB: 0.714\nCV: 0.822 LB: 0.616\nLooks like I overfit😂",
      "votes": 1,
      "replies": [
        {
          "id": 3070065,
          "postDate": "2024-12-12T07:38:14.617Z",
          "content": "<p>render the segmentation validation results for all training epoch. it may be obvious why the results are moving up and down</p>",
          "rawMarkdown": "render the segmentation validation results for all training epoch. it may be obvious why the results are moving up and down",
          "votes": 2,
          "replies": [
            {
              "id": 3070329,
              "postDate": "2024-12-12T15:05:34.410Z",
              "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Do you have any suggestions for visualization tools? I have tried plotly and matplotlib, but its hard to visualize the sample/predictions/labels simultaneously..</p>",
              "rawMarkdown": "@hengck23 Do you have any suggestions for visualization tools? I have tried plotly and matplotlib, but its hard to visualize the sample/predictions/labels simultaneously.."
            },
            {
              "id": 3070520,
              "postDate": "2024-12-12T19:23:24.827Z",
              "content": "<p>Not exactly the same thing, but would something likke this work?</p>\n<p><img src=\"https://www.kaggleusercontent.com/kf/211965594/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..87Gr-sO8M40yIaLaptjdzw.GtUH2o9w4EE_1-e2jAyMf06P6SS8fDfPc2e6GJZJyN7z5XbTbzyyGuWxo5dujwOMWG4g84zqIngbuY_xdIwi8ra5pClb-klx0Cen2IMTb8bnKGSbO4HvOfnY0KpQVsxsGX8_p7YIwFbUOpLQP1Yab2daVKvmsbulUN_6_ejJI8uroT2Bg023NDLcMagbo4JloAQvCGAJIFyEe2yNQPDakJDTU0RhHJiE7pVJ47Bp6XM1L2R1F0zOI6jtoxT3VxQ4uLPstgGhjixJ9cbW6U9D89pKzriRm4tXp-A0Jv-PytgBMNOrjlWUXKR6nlrBfG_2QITJgpm_uTI5jHj6IIfRA6fU30jkn00V_UIo23GOvMfCyEVx_mxYDolREAEoy_VtYn63JTtmljU7kWHK-K0etTMxFRiIN18dhjGNEFha8kl4U2IRM9kez--HmCTokAfeIRIkjTUOXopx9DD-qmngLMgQLQwgfDKn0bXAmEvCQ5JFgfKXXGY-_1qk66BgGOKoxeguTdoAWAFl7Ke0Vh8kOr1cOpEOz1hnuI8rquvEQXo4isniXXn0HfVfBCNlq7DtXC-YQFNn140jPEjcbbueeNz0IzMSt5nNnFCpAAU0eAU9HPee-LrZO4lfsWBZRlRrC1kBMxm-izKft9ALYMxooA.vMTUpush8MBOQ1mu2-WEdg/__results___files/__results___21_0.png\" alt=\"\"></p>\n<p>Code is here:<br>\n<a href=\"https://www.kaggle.com/code/davidlist/simulated-data-and-labels\" target=\"_blank\">https://www.kaggle.com/code/davidlist/simulated-data-and-labels</a></p>",
              "rawMarkdown": "Not exactly the same thing, but would something likke this work?\n\n![](https://www.kaggleusercontent.com/kf/211965594/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..87Gr-sO8M40yIaLaptjdzw.GtUH2o9w4EE_1-e2jAyMf06P6SS8fDfPc2e6GJZJyN7z5XbTbzyyGuWxo5dujwOMWG4g84zqIngbuY_xdIwi8ra5pClb-klx0Cen2IMTb8bnKGSbO4HvOfnY0KpQVsxsGX8_p7YIwFbUOpLQP1Yab2daVKvmsbulUN_6_ejJI8uroT2Bg023NDLcMagbo4JloAQvCGAJIFyEe2yNQPDakJDTU0RhHJiE7pVJ47Bp6XM1L2R1F0zOI6jtoxT3VxQ4uLPstgGhjixJ9cbW6U9D89pKzriRm4tXp-A0Jv-PytgBMNOrjlWUXKR6nlrBfG_2QITJgpm_uTI5jHj6IIfRA6fU30jkn00V_UIo23GOvMfCyEVx_mxYDolREAEoy_VtYn63JTtmljU7kWHK-K0etTMxFRiIN18dhjGNEFha8kl4U2IRM9kez--HmCTokAfeIRIkjTUOXopx9DD-qmngLMgQLQwgfDKn0bXAmEvCQ5JFgfKXXGY-_1qk66BgGOKoxeguTdoAWAFl7Ke0Vh8kOr1cOpEOz1hnuI8rquvEQXo4isniXXn0HfVfBCNlq7DtXC-YQFNn140jPEjcbbueeNz0IzMSt5nNnFCpAAU0eAU9HPee-LrZO4lfsWBZRlRrC1kBMxm-izKft9ALYMxooA.vMTUpush8MBOQ1mu2-WEdg/__results___files/__results___21_0.png)\n\nCode is here:\n[https://www.kaggle.com/code/davidlist/simulated-data-and-labels](https://www.kaggle.com/code/davidlist/simulated-data-and-labels)",
              "votes": 1
            },
            {
              "id": 3070541,
              "postDate": "2024-12-12T20:29:07.060Z",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/davidlist\" target=\"_blank\">@davidlist</a>, this is a nice 2D option! </p>\n<p>I was hoping for a 3D tool like the tomogram viewer on the CryoET Data Portal. It would be nice to be able to see all frames/labels/predictions and predictions in one visualization.</p>",
              "rawMarkdown": "Thanks @davidlist, this is a nice 2D option! \n\nI was hoping for a 3D tool like the tomogram viewer on the CryoET Data Portal. It would be nice to be able to see all frames/labels/predictions and predictions in one visualization."
            },
            {
              "id": 3070600,
              "postDate": "2024-12-12T22:00:09.407Z",
              "content": "<p>Ahh, right.  And for every epoch…</p>\n<p>This has <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> 's 3d viewer code at the end I believe:</p>\n<p><a href=\"https://www.kaggle.com/code/fnands/baseline-unet-train-submit\" target=\"_blank\">https://www.kaggle.com/code/fnands/baseline-unet-train-submit</a></p>",
              "rawMarkdown": "Ahh, right.  And for every epoch...\n\nThis has @hengck23 's 3d viewer code at the end I believe:\n\n[https://www.kaggle.com/code/fnands/baseline-unet-train-submit](https://www.kaggle.com/code/fnands/baseline-unet-train-submit)",
              "votes": 1
            },
            {
              "id": 3073776,
              "postDate": "2024-12-16T22:41:39.090Z",
              "content": "<p>Found a great tool called <a href=\"https://napari.org/stable/\" target=\"_blank\">Napari</a>. I dont  think it will run in the Kaggle environment, but great for local visualization. Would highly recommend!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F89195b6071b126a39bf569f2edc13a05%2FCapture.JPG?generation=1734389124881687&amp;alt=media\" alt=\"img\"></p>",
              "rawMarkdown": "Found a great tool called [Napari](https://napari.org/stable/). I dont  think it will run in the Kaggle environment, but great for local visualization. Would highly recommend!\n\n![img](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F89195b6071b126a39bf569f2edc13a05%2FCapture.JPG?generation=1734389124881687&alt=media)",
              "votes": 3
            }
          ]
        }
      ]
    },
    {
      "id": 3065442,
      "postDate": "2024-12-06T19:46:47.327Z",
      "content": "<p>I train with all samples and seven folds. For CV I predict each sample with his out of bag model. I join all the predicitons in one \"submission\" and I apply metric over it. </p>\n<p>CV           .728 .713 .739 .745 .777</p>\n<p>LB             .678 .649 .693 .687 .716</p>\n<p>LB Refined  .702 .685 .703 .711 .718</p>\n<p></p>\n<p></p>\n<p>EDIT: Bug fixed, when collecting GT I've been using \"target\" variable that was not changing in the loop. So I've been collecting virus coordinates 5 times per sample. Now everything runs as espected. I'll update my results as soon as possible. \"Fortunately\" there is still about two months.</p>",
      "rawMarkdown": "I train with all samples and seven folds. For CV I predict each sample with his out of bag model. I join all the predicitons in one \"submission\" and I apply metric over it. ~~The CV is consitent with LB. But is worst than LB, what is opposite to what people are reporting.~~\n\nCV ~~ .532 .554~~          .728 .713 .739 .745 .777\n\nLB ~~.624 .678~~            .678 .649 .693 .687 .716\n\nLB Refined ~~.??? .???~~ .702 .685 .703 .711 .718\n\n~~And that's with intern metric radius in A but predictions and GT in 10A (what applyes a radius 10 times bigger than it should when catching positives). With predictions and GT in A CV decreases to ~.1 (all of this ith distance_multiplier = .5 and beta = 4).~~\n\n~~I really don't know why, but I'll continue this way since show the most coherent results for me.~~\n\nEDIT: Bug fixed, when collecting GT I've been using \"target\" variable that was not changing in the loop. So I've been collecting virus coordinates 5 times per sample. Now everything runs as espected. I'll update my results as soon as possible. \"Fortunately\" there is still about two months.",
      "votes": 1
    },
    {
      "id": 3065127,
      "postDate": "2024-12-06T11:37:39.637Z",
      "content": "<p>Another question: which GT do you use to compute your metric? Is it the point coordinate extracted from blobs or the original coordinate from .jsons ?</p>\n<p>We have only started so it is just a matter of 1 exp:<br>\nCV: 0.71452<br>\nLB: 0.692</p>",
      "rawMarkdown": "Another question: which GT do you use to compute your metric? Is it the point coordinate extracted from blobs or the original coordinate from .jsons ?\n\nWe have only started so it is just a matter of 1 exp:\nCV: 0.71452\nLB: 0.692",
      "votes": 1,
      "replies": [
        {
          "id": 3065177,
          "postDate": "2024-12-06T13:38:48.437Z",
          "content": "<p>GT - originals<br>\nPredicted - extracted from blobs</p>",
          "rawMarkdown": "GT - originals\nPredicted - extracted from blobs",
          "votes": 1
        }
      ]
    },
    {
      "id": 3063685,
      "postDate": "2024-12-04T18:54:34.683Z",
      "content": "<p>CV 0.734<br>\nLB 0.639</p>",
      "rawMarkdown": "CV 0.734\nLB 0.639",
      "votes": 1
    },
    {
      "id": 3061598,
      "postDate": "2024-12-02T20:48:50.973Z",
      "content": "<p>CV 0.718, LB 0.618. I'm pretty consistently getting a gap of ~0.1 if I don't do anything weird with my models.</p>",
      "rawMarkdown": "CV 0.718, LB 0.618. I'm pretty consistently getting a gap of ~0.1 if I don't do anything weird with my models.",
      "votes": 1
    },
    {
      "id": 3066758,
      "postDate": "2024-12-08T12:52:36.193Z",
      "content": "<p>Do you guys calculate the competition metric each epoch? I do, and I can get anywhere between 0.5 and 0.78-0.79 peaks as my local LB (I use 6 train files and 1 validation file). My highest on the leaderboard was achieved submitting 0.757 local LB model vs. 0.667 on the public LB.</p>",
      "rawMarkdown": "Do you guys calculate the competition metric each epoch? I do, and I can get anywhere between 0.5 and 0.78-0.79 peaks as my local LB (I use 6 train files and 1 validation file). My highest on the leaderboard was achieved submitting 0.757 local LB model vs. 0.667 on the public LB.",
      "votes": 2,
      "replies": [
        {
          "id": 3069062,
          "postDate": "2024-12-11T02:43:03.417Z",
          "content": "<p>That's probably the best way to do it.</p>",
          "rawMarkdown": "That's probably the best way to do it."
        },
        {
          "id": 3069882,
          "postDate": "2024-12-12T02:05:22.283Z",
          "content": "<p>Hi Andrei, </p>\n<p>May I ask if we could find a ready-to-use competition metric code demo? I just notice there is a description of the competition metric (F-Beta with a beta of 4). </p>\n<p>Best <br>\nLeo</p>",
          "rawMarkdown": "Hi Andrei, \n\nMay I ask if we could find a ready-to-use competition metric code demo? I just notice there is a description of the competition metric (F-Beta with a beta of 4). \n\nBest \nLeo",
          "replies": [
            {
              "id": 3070166,
              "postDate": "2024-12-12T11:09:12.510Z",
              "content": "<p>Hello Leo, there's an official competition metric provided by the host <a href=\"https://www.kaggle.com/code/metric/czi-cryoet-84969\" target=\"_blank\">here</a>. I'll be honest, I just used <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>'s pipeline in <a href=\"https://www.kaggle.com/code/hengck23/3d-unet-using-2d-image-encoder\" target=\"_blank\">one</a> of his public codes. The most common framework used that I see is softmaxing the logit output into probabilities and making them hard labels based on a threshold condition, followed by using cc3d to label the objects and derive centroids. </p>\n<p>Then, you create your submission.csv based on the format required by the host (id,experiment,particle_type,x,y,z) and that gets scored against the label centroid location of the particles, where a centroid is considered a true positive as long as it matches the experiment, particle type and is within 0.5 radius of the particle type. You compute FP, FN, TP, TN and calculate the fbeta.</p>",
              "rawMarkdown": "Hello Leo, there's an official competition metric provided by the host [here](https://www.kaggle.com/code/metric/czi-cryoet-84969). I'll be honest, I just used @hengck23's pipeline in [one](https://www.kaggle.com/code/hengck23/3d-unet-using-2d-image-encoder) of his public codes. The most common framework used that I see is softmaxing the logit output into probabilities and making them hard labels based on a threshold condition, followed by using cc3d to label the objects and derive centroids. \n\nThen, you create your submission.csv based on the format required by the host (id,experiment,particle_type,x,y,z) and that gets scored against the label centroid location of the particles, where a centroid is considered a true positive as long as it matches the experiment, particle type and is within 0.5 radius of the particle type. You compute FP, FN, TP, TN and calculate the fbeta."
            },
            {
              "id": 3072468,
              "postDate": "2024-12-15T08:50:07.607Z",
              "content": "<p>thank you so much!!</p>",
              "rawMarkdown": "thank you so much!!"
            }
          ]
        }
      ]
    },
    {
      "id": 3066518,
      "postDate": "2024-12-08T08:47:30.973Z",
      "content": "<p>So…  Mine was off because I was calculating my CV from the three test files, two of which I trained against.  When I calculate my CV from only my local hold out, I get 0.66 vs 0.65 for my LB score.</p>",
      "rawMarkdown": "So...  Mine was off because I was calculating my CV from the three test files, two of which I trained against.  When I calculate my CV from only my local hold out, I get 0.66 vs 0.65 for my LB score.",
      "votes": 2,
      "replies": [
        {
          "id": 3066535,
          "postDate": "2024-12-08T08:59:42.717Z",
          "content": "<p>I'm glad you've found the culprit. I've just submitted 4 folds. Will see how it goes.</p>",
          "rawMarkdown": "I'm glad you've found the culprit. I've just submitted 4 folds. Will see how it goes.",
          "votes": 1,
          "replies": [
            {
              "id": 3067089,
              "postDate": "2024-12-08T21:28:09.837Z",
              "content": "<p>good job!  </p>",
              "rawMarkdown": "good job!  ",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 3061684,
      "postDate": "2024-12-02T23:35:30.017Z",
      "content": "<p>CV 0.73, LB 0.61.</p>",
      "rawMarkdown": "CV 0.73, LB 0.61.",
      "votes": 2
    },
    {
      "id": 3081870,
      "postDate": "2024-12-27T10:26:42.280Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3061768,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-12-03T01:14:09.827000",
      "content": "<p>please imagine how the loss curves for train,validation and hidden test would look like over different numbers of train, validation and test samples.<br>\n(Hint: if you cannot imagine how the curve look like, download cifar10 and do experiments and plot the curve for different umbers of train,validation and  test. assumption here is that the data difficulty of cirfar10 is same as the kaggle competition, if not \"imagine\" the change in cifar curves to be accounted for when \"transfering them\" to this kaggle dataset)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc27daac31cf753118576333f50fe3ec1%2FSelection_198.png?generation=1733224281913535&amp;alt=media\" alt=\"\"></p>\n<p>there should be 4 curve (train,valid, hidden public/private test). i only show two (train, valid). you can probe for the rest.</p>\n<p>you are reporting only one black point on the curve. It is better to consider the whole curve itself (i.e. submitting over multiple models at different epochs)</p>\n<hr>\n<p>once you understand this concept, then \"shakeup or not\", \"lb/cv correlation\", etc is only statistics.</p>\n<p>here \"parameters\" are : num of train, num of valid samples, model complexity and we are trying to </p>\n<ol>\n<li>make hidden test curve overlaping with our valid curve as much as possible. </li>\n<li>select a point (i.e. how many epoch to train model) in the overlapping region for submission</li>\n</ol>",
      "votes": 14,
      "replies": [
        {
          "id": 3061957,
          "author_name": "Timmy Juicehouse",
          "author_url": "",
          "post_date": "2024-12-03T06:46:50.103000",
          "content": "<p>Great thanks, your post resolves my confusion in this competition.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3061966,
          "author_name": "DennisSakva",
          "author_url": "",
          "post_date": "2024-12-03T07:06:44.703000",
          "content": "<p>Hi, Heng. Thanks for the explanation. I wasn’t looking for the theoretical reasons behind the difference. I just wanted to gather a few data points on what others are observing regarding the CV-LB discrepancy. From the few I’ve seen so far, the difference seems to consistently be around 0.1. Three data points aren’t enough for any definitive conclusions, but it’s at least something to think about! <br>\nCheers.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3061975,
              "author_name": "David List",
              "author_url": "",
              "post_date": "2024-12-03T07:19:04.377000",
              "content": "<p>It could also be that we're all using the same tomogram for validation.  😀</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3061983,
              "author_name": "DennisSakva",
              "author_url": "",
              "post_date": "2024-12-03T07:31:48.907000",
              "content": "<p>Sure thing! But consistency is a good sign and allows to infer which files give better estimates of which particles.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3061992,
              "author_name": "David List",
              "author_url": "",
              "post_date": "2024-12-03T07:39:12.490000",
              "content": "<p>TS_6_4, which I think a lot of people are using may be the odd one of the 7, because it's the only one with the strange viruses.  I doubt that explains everything, but it might be a small piece…  Actually, scratch that.  If TS_6_4 is more like the LB data, that would have the opposite effect.<br>\n<img src=\"https://www.kaggleusercontent.com/kf/210412070/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..RNYcK5DUtlQkP555-N8opg.lnLUsGlsB9Ri3l5VCUUlLYTJ0NKlPmh90c2dO_ua-vvzPdjsvHfjDGBJT-be5BrHutGtL0ftq-o0xHLcqlfAQhr7Y38sjkrOh372vggex2nEG7Fq8O4UiRyaD-7v6_MnJHlkaSq_heDXXDGFsQjJFUbyafwDxIh_bSwJZLQrBzfNBDoBkJOBdpXKJi2hFifZxIlsoyXvLv8rheoUXC6lrQa8xl1klTL3DT71p_bYBGwR7l2_8e9ENFFEVOWA7uRAEdoc6Bo6A0KKA2ZjGCXzNcYNYjE4Vsa8c_65W48Ly7IUbeAboGd1B6f0LeKwzM4TN3NqG2jCOKRa2IsF_Y-6Idl9PhfaOP3_0LHp0SnU0Jzs4QTgD9Gi0TrnfNIXqjYuna9BsHR6KwQPZgc8B_8emlDLywb3UM3YsOXaLAUA8A9c637jc1tjZLMlVtPqWtHOn8zbpqIrNX7OynQZOGZcQeuBE1pIvcF51A8FT-PSCb4-bjknY_x2e142wHZCQiVhHIfFI0HDftNLoC6jjfFSZaMnFtZhaAeWWLGoVknSTDXNH189QObu-xwouffCB3QK84Sc87SCnrivKaaeA9U6ndshiwrjPCxdYU2RZKYsn4T4SiY7HHw4wXE-esPJEuLK551R36c4Z5bW-nsEaOm8jg.CyN4EmtJly4lEqLASAA8rA/__results___files/__results___9_0.png\" alt=\"\"></p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3062004,
              "author_name": "David List",
              "author_url": "",
              "post_date": "2024-12-03T07:52:37",
              "content": "<p>Sorry, that image is a bit longer than I remembered.  😀</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3062106,
              "author_name": "DennisSakva",
              "author_url": "",
              "post_date": "2024-12-03T09:48:34.577000",
              "content": "<p>But it's cool :)</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3070925,
      "author_name": "Tanuki_boosting",
      "author_url": "",
      "post_date": "2024-12-13T07:06:25.123000",
      "content": "<p>The way Cross Validation is split affects the gap between CV and LB scores. However, up to a certain score threshold, it seems that CV and LB are correlated.<br>\nEssentially, I believe this phenomenon occurs when the training data includes samples similar to those used for calculating the Public Leaderboard, resulting in higher LB scores.</p>\n<p>CV0.77 LB0.7</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3070012,
      "author_name": "MOONMOON",
      "author_url": "",
      "post_date": "2024-12-12T06:15:16.947000",
      "content": "<p>Validation on TS_6_4<br>\nCV: 0.755 LB: 0.714<br>\nCV: 0.822 LB: 0.616<br>\nLooks like I overfit😂</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3070065,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2024-12-12T07:38:14.617000",
          "content": "<p>render the segmentation validation results for all training epoch. it may be obvious why the results are moving up and down</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3070329,
              "author_name": "Bartley",
              "author_url": "",
              "post_date": "2024-12-12T15:05:34.410000",
              "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Do you have any suggestions for visualization tools? I have tried plotly and matplotlib, but its hard to visualize the sample/predictions/labels simultaneously..</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3070520,
              "author_name": "David List",
              "author_url": "",
              "post_date": "2024-12-12T19:23:24.827000",
              "content": "<p>Not exactly the same thing, but would something likke this work?</p>\n<p><img src=\"https://www.kaggleusercontent.com/kf/211965594/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..87Gr-sO8M40yIaLaptjdzw.GtUH2o9w4EE_1-e2jAyMf06P6SS8fDfPc2e6GJZJyN7z5XbTbzyyGuWxo5dujwOMWG4g84zqIngbuY_xdIwi8ra5pClb-klx0Cen2IMTb8bnKGSbO4HvOfnY0KpQVsxsGX8_p7YIwFbUOpLQP1Yab2daVKvmsbulUN_6_ejJI8uroT2Bg023NDLcMagbo4JloAQvCGAJIFyEe2yNQPDakJDTU0RhHJiE7pVJ47Bp6XM1L2R1F0zOI6jtoxT3VxQ4uLPstgGhjixJ9cbW6U9D89pKzriRm4tXp-A0Jv-PytgBMNOrjlWUXKR6nlrBfG_2QITJgpm_uTI5jHj6IIfRA6fU30jkn00V_UIo23GOvMfCyEVx_mxYDolREAEoy_VtYn63JTtmljU7kWHK-K0etTMxFRiIN18dhjGNEFha8kl4U2IRM9kez--HmCTokAfeIRIkjTUOXopx9DD-qmngLMgQLQwgfDKn0bXAmEvCQ5JFgfKXXGY-_1qk66BgGOKoxeguTdoAWAFl7Ke0Vh8kOr1cOpEOz1hnuI8rquvEQXo4isniXXn0HfVfBCNlq7DtXC-YQFNn140jPEjcbbueeNz0IzMSt5nNnFCpAAU0eAU9HPee-LrZO4lfsWBZRlRrC1kBMxm-izKft9ALYMxooA.vMTUpush8MBOQ1mu2-WEdg/__results___files/__results___21_0.png\" alt=\"\"></p>\n<p>Code is here:<br>\n<a href=\"https://www.kaggle.com/code/davidlist/simulated-data-and-labels\" target=\"_blank\">https://www.kaggle.com/code/davidlist/simulated-data-and-labels</a></p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3070541,
              "author_name": "Bartley",
              "author_url": "",
              "post_date": "2024-12-12T20:29:07.060000",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/davidlist\" target=\"_blank\">@davidlist</a>, this is a nice 2D option! </p>\n<p>I was hoping for a 3D tool like the tomogram viewer on the CryoET Data Portal. It would be nice to be able to see all frames/labels/predictions and predictions in one visualization.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3070600,
              "author_name": "David List",
              "author_url": "",
              "post_date": "2024-12-12T22:00:09.407000",
              "content": "<p>Ahh, right.  And for every epoch…</p>\n<p>This has <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> 's 3d viewer code at the end I believe:</p>\n<p><a href=\"https://www.kaggle.com/code/fnands/baseline-unet-train-submit\" target=\"_blank\">https://www.kaggle.com/code/fnands/baseline-unet-train-submit</a></p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3073776,
              "author_name": "Bartley",
              "author_url": "",
              "post_date": "2024-12-16T22:41:39.090000",
              "content": "<p>Found a great tool called <a href=\"https://napari.org/stable/\" target=\"_blank\">Napari</a>. I dont  think it will run in the Kaggle environment, but great for local visualization. Would highly recommend!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F89195b6071b126a39bf569f2edc13a05%2FCapture.JPG?generation=1734389124881687&amp;alt=media\" alt=\"img\"></p>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3065442,
      "author_name": "Ángel Jacinto Sánchez Ruiz",
      "author_url": "",
      "post_date": "2024-12-06T19:46:47.327000",
      "content": "<p>I train with all samples and seven folds. For CV I predict each sample with his out of bag model. I join all the predicitons in one \"submission\" and I apply metric over it. </p>\n<p>CV           .728 .713 .739 .745 .777</p>\n<p>LB             .678 .649 .693 .687 .716</p>\n<p>LB Refined  .702 .685 .703 .711 .718</p>\n<p></p>\n<p></p>\n<p>EDIT: Bug fixed, when collecting GT I've been using \"target\" variable that was not changing in the loop. So I've been collecting virus coordinates 5 times per sample. Now everything runs as espected. I'll update my results as soon as possible. \"Fortunately\" there is still about two months.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3065127,
      "author_name": "Volodymyr",
      "author_url": "",
      "post_date": "2024-12-06T11:37:39.637000",
      "content": "<p>Another question: which GT do you use to compute your metric? Is it the point coordinate extracted from blobs or the original coordinate from .jsons ?</p>\n<p>We have only started so it is just a matter of 1 exp:<br>\nCV: 0.71452<br>\nLB: 0.692</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3065177,
          "author_name": "DennisSakva",
          "author_url": "",
          "post_date": "2024-12-06T13:38:48.437000",
          "content": "<p>GT - originals<br>\nPredicted - extracted from blobs</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3063685,
      "author_name": "Sergio Alvarez",
      "author_url": "",
      "post_date": "2024-12-04T18:54:34.683000",
      "content": "<p>CV 0.734<br>\nLB 0.639</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3061598,
      "author_name": "Jeroen Cottaar",
      "author_url": "",
      "post_date": "2024-12-02T20:48:50.973000",
      "content": "<p>CV 0.718, LB 0.618. I'm pretty consistently getting a gap of ~0.1 if I don't do anything weird with my models.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3066758,
      "author_name": "Andrei Zamfir",
      "author_url": "",
      "post_date": "2024-12-08T12:52:36.193000",
      "content": "<p>Do you guys calculate the competition metric each epoch? I do, and I can get anywhere between 0.5 and 0.78-0.79 peaks as my local LB (I use 6 train files and 1 validation file). My highest on the leaderboard was achieved submitting 0.757 local LB model vs. 0.667 on the public LB.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3069062,
          "author_name": "David List",
          "author_url": "",
          "post_date": "2024-12-11T02:43:03.417000",
          "content": "<p>That's probably the best way to do it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3069882,
          "author_name": "Leo Yang",
          "author_url": "",
          "post_date": "2024-12-12T02:05:22.283000",
          "content": "<p>Hi Andrei, </p>\n<p>May I ask if we could find a ready-to-use competition metric code demo? I just notice there is a description of the competition metric (F-Beta with a beta of 4). </p>\n<p>Best <br>\nLeo</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3070166,
              "author_name": "Andrei Zamfir",
              "author_url": "",
              "post_date": "2024-12-12T11:09:12.510000",
              "content": "<p>Hello Leo, there's an official competition metric provided by the host <a href=\"https://www.kaggle.com/code/metric/czi-cryoet-84969\" target=\"_blank\">here</a>. I'll be honest, I just used <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>'s pipeline in <a href=\"https://www.kaggle.com/code/hengck23/3d-unet-using-2d-image-encoder\" target=\"_blank\">one</a> of his public codes. The most common framework used that I see is softmaxing the logit output into probabilities and making them hard labels based on a threshold condition, followed by using cc3d to label the objects and derive centroids. </p>\n<p>Then, you create your submission.csv based on the format required by the host (id,experiment,particle_type,x,y,z) and that gets scored against the label centroid location of the particles, where a centroid is considered a true positive as long as it matches the experiment, particle type and is within 0.5 radius of the particle type. You compute FP, FN, TP, TN and calculate the fbeta.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3072468,
              "author_name": "Leo Yang",
              "author_url": "",
              "post_date": "2024-12-15T08:50:07.607000",
              "content": "<p>thank you so much!!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3066518,
      "author_name": "David List",
      "author_url": "",
      "post_date": "2024-12-08T08:47:30.973000",
      "content": "<p>So…  Mine was off because I was calculating my CV from the three test files, two of which I trained against.  When I calculate my CV from only my local hold out, I get 0.66 vs 0.65 for my LB score.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3066535,
          "author_name": "DennisSakva",
          "author_url": "",
          "post_date": "2024-12-08T08:59:42.717000",
          "content": "<p>I'm glad you've found the culprit. I've just submitted 4 folds. Will see how it goes.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3067089,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-12-08T21:28:09.837000",
              "content": "<p>good job!  </p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3061684,
      "author_name": "David List",
      "author_url": "",
      "post_date": "2024-12-02T23:35:30.017000",
      "content": "<p>CV 0.73, LB 0.61.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3081870,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-12-27T10:26:42.280000",
      "content": "",
      "votes": -1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3061768": "please imagine how the loss curves for train,validation and hidden test would look like over different numbers of train, validation and test samples.\n(Hint: if you cannot imagine how the curve look like, download cifar10 and do experiments and plot the curve for different umbers of train,validation and  test. assumption here is that the data difficulty of cirfar10 is same as the kaggle competition, if not \"imagine\" the change in cifar curves to be accounted for when \"transfering them\" to this kaggle dataset)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc27daac31cf753118576333f50fe3ec1%2FSelection_198.png?generation=1733224281913535&alt=media)\n\n\nthere should be 4 curve (train,valid, hidden public/private test). i only show two (train, valid). you can probe for the rest.\n\nyou are reporting only one black point on the curve. It is better to consider the whole curve itself (i.e. submitting over multiple models at different epochs)\n\n---\n\nonce you understand this concept, then \"shakeup or not\", \"lb/cv correlation\", etc is only statistics.\n\nhere \"parameters\" are : num of train, num of valid samples, model complexity and we are trying to \n1. make hidden test curve overlaping with our valid curve as much as possible. \n2. select a point (i.e. how many epoch to train model) in the overlapping region for submission",
    "3061477": "I'll start. Train - first five, validation on the last two\nCV 0.8 LB 0.7 with poor correlation between CV and LB",
    "3070925": "The way Cross Validation is split affects the gap between CV and LB scores. However, up to a certain score threshold, it seems that CV and LB are correlated.\nEssentially, I believe this phenomenon occurs when the training data includes samples similar to those used for calculating the Public Leaderboard, resulting in higher LB scores.\n\nCV0.77 LB0.7",
    "3070012": "Validation on TS_6_4\nCV: 0.755 LB: 0.714\nCV: 0.822 LB: 0.616\nLooks like I overfit😂",
    "3065442": "I train with all samples and seven folds. For CV I predict each sample with his out of bag model. I join all the predicitons in one \"submission\" and I apply metric over it. ~~The CV is consitent with LB. But is worst than LB, what is opposite to what people are reporting.~~\n\nCV ~~ .532 .554~~          .728 .713 .739 .745 .777\n\nLB ~~.624 .678~~            .678 .649 .693 .687 .716\n\nLB Refined ~~.??? .???~~ .702 .685 .703 .711 .718\n\n~~And that's with intern metric radius in A but predictions and GT in 10A (what applyes a radius 10 times bigger than it should when catching positives). With predictions and GT in A CV decreases to ~.1 (all of this ith distance_multiplier = .5 and beta = 4).~~\n\n~~I really don't know why, but I'll continue this way since show the most coherent results for me.~~\n\nEDIT: Bug fixed, when collecting GT I've been using \"target\" variable that was not changing in the loop. So I've been collecting virus coordinates 5 times per sample. Now everything runs as espected. I'll update my results as soon as possible. \"Fortunately\" there is still about two months.",
    "3065127": "Another question: which GT do you use to compute your metric? Is it the point coordinate extracted from blobs or the original coordinate from .jsons ?\n\nWe have only started so it is just a matter of 1 exp:\nCV: 0.71452\nLB: 0.692",
    "3063685": "CV 0.734\nLB 0.639",
    "3061598": "CV 0.718, LB 0.618. I'm pretty consistently getting a gap of ~0.1 if I don't do anything weird with my models.",
    "3066758": "Do you guys calculate the competition metric each epoch? I do, and I can get anywhere between 0.5 and 0.78-0.79 peaks as my local LB (I use 6 train files and 1 validation file). My highest on the leaderboard was achieved submitting 0.757 local LB model vs. 0.667 on the public LB.",
    "3066518": "So...  Mine was off because I was calculating my CV from the three test files, two of which I trained against.  When I calculate my CV from only my local hold out, I get 0.66 vs 0.65 for my LB score.",
    "3061684": "CV 0.73, LB 0.61.",
    "3081870": ""
  }
}