{
  "id": 550674,
  "title": "Ensemble Tips?",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/550674",
  "author_name": "",
  "post_date": "2024-12-08T22:24:48.392846200Z",
  "votes": 4,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Hey everyone,</p>\n<p>I'm currently working with single fold models and I'd like to find a way to ensemble my cross-validation models. My current train F_beta is 0.76, validation (two samples) is 0.75, and my LB score with this single fold model is 0.639. My hope is that building a model that's an ensemble of the cross-validation scores will help me reduce the gap between the leaderboard and my validation data. Note: it's entirely possible that this LB score is closer to my true cross-validation score and perhaps I'm validating on two 'easy' samples and I'm working on validating this.</p>\n<p>Of course I'm going to experiment with this on my own, but I'm looking to learn from the Kaggle community's experience here. When trying to keep your submissions &lt; 12 hours, how do you usually approach this? I've done some reading and used Chat GPT to get me started. Here are 4 things I think sound reasonable and I'd like to hear your opinions/experiences.</p>\n<p>1) Average predictions from each of the folded models independently. This seems straightforward and I've used it before, but it could take some time. I've got some headroom (my current process takes about 6 hours to score), but I'd like to learn about other more efficient ways to do this.</p>\n<p>2) Weight interpolation. This seems interesting, performant, and straight-forward, but it's a bit of a new concept for me. Have you had success with this in the past? It seems logical, but it also seems like things could go wrong very quickly.</p>\n<p>3) Distillation. This seems like a neat way to get one smaller model that is a blend of the many, but I'm a little worried that it would be prone to overfitting since having a true train/test split here would be tricky with the number of samples we have. Since the models we're distilling from are from different folds, it's hard to generate a true holdout set here with only 7 samples.</p>\n<p>4) Use CV to identify the best learning rate schedule, then apply that schedule to all 7 samples. This sounds pretty easy, but best practices about how to validate this sort of thing (outside of trusting the leaderboard) seems kind of tough. There's a little more 'faith' behind this approach than I'd like to have. </p>",
  "messages": [
    {
      "id": "3067109",
      "postDate": "12/08/2024 22:24:48",
      "content": "<p>Hey everyone,</p>\n<p>I'm currently working with single fold models and I'd like to find a way to ensemble my cross-validation models. My current train F_beta is 0.76, validation (two samples) is 0.75, and my LB score with this single fold model is 0.639. My hope is that building a model that's an ensemble of the cross-validation scores will help me reduce the gap between the leaderboard and my validation data. Note: it's entirely possible that this LB score is closer to my true cross-validation score and perhaps I'm validating on two 'easy' samples and I'm working on validating this.</p>\n<p>Of course I'm going to experiment with this on my own, but I'm looking to learn from the Kaggle community's experience here. When trying to keep your submissions &lt; 12 hours, how do you usually approach this? I've done some reading and used Chat GPT to get me started. Here are 4 things I think sound reasonable and I'd like to hear your opinions/experiences.</p>\n<p>1) Average predictions from each of the folded models independently. This seems straightforward and I've used it before, but it could take some time. I've got some headroom (my current process takes about 6 hours to score), but I'd like to learn about other more efficient ways to do this.</p>\n<p>2) Weight interpolation. This seems interesting, performant, and straight-forward, but it's a bit of a new concept for me. Have you had success with this in the past? It seems logical, but it also seems like things could go wrong very quickly.</p>\n<p>3) Distillation. This seems like a neat way to get one smaller model that is a blend of the many, but I'm a little worried that it would be prone to overfitting since having a true train/test split here would be tricky with the number of samples we have. Since the models we're distilling from are from different folds, it's hard to generate a true holdout set here with only 7 samples.</p>\n<p>4) Use CV to identify the best learning rate schedule, then apply that schedule to all 7 samples. This sounds pretty easy, but best practices about how to validate this sort of thing (outside of trusting the leaderboard) seems kind of tough. There's a little more 'faith' behind this approach than I'd like to have. </p>",
      "rawMarkdown": "Hey everyone,\n\nI'm currently working with single fold models and I'd like to find a way to ensemble my cross-validation models. My current train F_beta is 0.76, validation (two samples) is 0.75, and my LB score with this single fold model is 0.639. My hope is that building a model that's an ensemble of the cross-validation scores will help me reduce the gap between the leaderboard and my validation data. Note: it's entirely possible that this LB score is closer to my true cross-validation score and perhaps I'm validating on two 'easy' samples and I'm working on validating this.\n\nOf course I'm going to experiment with this on my own, but I'm looking to learn from the Kaggle community's experience here. When trying to keep your submissions < 12 hours, how do you usually approach this? I've done some reading and used Chat GPT to get me started. Here are 4 things I think sound reasonable and I'd like to hear your opinions/experiences.\n\n1) Average predictions from each of the folded models independently. This seems straightforward and I've used it before, but it could take some time. I've got some headroom (my current process takes about 6 hours to score), but I'd like to learn about other more efficient ways to do this.\n\n2) Weight interpolation. This seems interesting, performant, and straight-forward, but it's a bit of a new concept for me. Have you had success with this in the past? It seems logical, but it also seems like things could go wrong very quickly.\n\n3) Distillation. This seems like a neat way to get one smaller model that is a blend of the many, but I'm a little worried that it would be prone to overfitting since having a true train/test split here would be tricky with the number of samples we have. Since the models we're distilling from are from different folds, it's hard to generate a true holdout set here with only 7 samples.\n\n4) Use CV to identify the best learning rate schedule, then apply that schedule to all 7 samples. This sounds pretty easy, but best practices about how to validate this sort of thing (outside of trusting the leaderboard) seems kind of tough. There's a little more 'faith' behind this approach than I'd like to have.",
      "votes": null
    },
    {
      "id": "3067380",
      "postDate": "12/09/2024 08:01:09",
      "content": "<p>the first thing is to speed up your single model.<br>\ncheck the public code and post.<br>\nyou should be able to get about lb0.650 within 1.5 to 2 hours.</p>\n<p>distillation, etc. is only used as the last resort.</p>\n<p>if you cannot ensemble (and TTA) your score will not increase.<br>\nyou need these to get lb0.750 and beyond</p>",
      "rawMarkdown": "the first thing is to speed up your single model.\ncheck the public code and post.\nyou should be able to get about lb0.650 within 1.5 to 2 hours.\n\ndistillation, etc. is only used as the last resort.\n\nif you cannot ensemble (and TTA) your score will not increase.\nyou need these to get lb0.750 and beyond",
      "votes": null
    },
    {
      "id": "3067564",
      "postDate": "12/09/2024 12:11:57",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> ! I saw your pytorch code to speed up the connected components. I will spend some time to understand and incorporate that into my workflow. I think I can probably speed up how I generate the tiles too. This will give me more headroom for ensembling &amp; TTA. I'll try to find ways to make the model faster too.</p>",
      "rawMarkdown": "Thank you @hengck23 ! I saw your pytorch code to speed up the connected components. I will spend some time to understand and incorporate that into my workflow. I think I can probably speed up how I generate the tiles too. This will give me more headroom for ensembling & TTA. I'll try to find ways to make the model faster too.",
      "votes": null
    },
    {
      "id": "3070229",
      "postDate": "12/12/2024 12:58:11",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> you were right, ensembling can help the public score a lot! Finding the right thresholds helps a lot too. I still need to make the code faster so I can do TTA 😂. I've got a 7-fold ensemble running in ~11.5 hours haha…maybe 7 folds was a bit too much.</p>",
      "rawMarkdown": "hengck23 you were right, ensembling can help the public score a lot! Finding the right thresholds helps a lot too. I still need to make the code faster so I can do TTA 😂. I've got a 7-fold ensemble running in ~11.5 hours haha...maybe 7 folds was a bit too much.",
      "votes": null
    },
    {
      "id": "3070268",
      "postDate": "12/12/2024 13:44:05",
      "content": "<p>How much did ensembling improve your score?</p>",
      "rawMarkdown": "How much did ensembling improve your score?",
      "votes": null
    },
    {
      "id": "3070333",
      "postDate": "12/12/2024 15:13:07",
      "content": "<p>Still doing some validation on that by checking the scores from the individual models with optimized thresholds, but it looks like it will be by about 0.05 or so.</p>",
      "rawMarkdown": "Still doing some validation on that by checking the scores from the individual models with optimized thresholds, but it looks like it will be by about 0.05 or so.",
      "votes": null
    },
    {
      "id": "3070339",
      "postDate": "12/12/2024 15:17:44",
      "content": "<p>'much' more than that is possible. it really depends on your model accuracy, tta, etc.</p>\n<p>if you wan to know the theoretical limit, google or ask chatgpt.<br>\ntrain accuracy is a weak upper bound</p>",
      "rawMarkdown": "'much' more than that is possible. it really depends on your model accuracy, tta, etc.\n\nif you wan to know the theoretical limit, google or ask chatgpt.\ntrain accuracy is a weak upper bound",
      "votes": null
    },
    {
      "id": "3070340",
      "postDate": "12/12/2024 15:17:48",
      "content": "<p>Thanks, I haven't done ensembling yet as I ran out of GPU for the week, but with TTA I saw an increase of ~0.03.</p>",
      "rawMarkdown": "Thanks, I haven't done ensembling yet as I ran out of GPU for the week, but with TTA I saw an increase of ~0.03.",
      "votes": null
    },
    {
      "id": "3070453",
      "postDate": "12/12/2024 17:14:10",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> ! I believe that's absolutely possible. My current models are likely a bit overfit. I'm not using much augmentation yet. This ensemble was also trained with a fixed learning rate schedule I optimized for one of the folds and I actually think on some of the folds, this caused quite a bit of overfitting. There's still a pretty large gap (about 0.1) from my CV F_Beta score and my LB F_Beta score. I will keep trying!</p>",
      "rawMarkdown": "Thanks @hengck23 ! I believe that's absolutely possible. My current models are likely a bit overfit. I'm not using much augmentation yet. This ensemble was also trained with a fixed learning rate schedule I optimized for one of the folds and I actually think on some of the folds, this caused quite a bit of overfitting. There's still a pretty large gap (about 0.1) from my CV F_Beta score and my LB F_Beta score. I will keep trying!",
      "votes": null
    },
    {
      "id": "3070456",
      "postDate": "12/12/2024 17:16:01",
      "content": "<p><a href=\"https://www.kaggle.com/andreizamfir\" target=\"_blank\">@andreizamfir</a> good to know! That's pretty close to what I was hoping for/expecting. If I increase the tile overlaps for some of my smaller models, kind of like a small TTA in some ways, I see noticable improvement (about 0.01-0.02) so I suspect that proper TTA will be a larger boost.</p>",
      "rawMarkdown": "andreizamfir good to know! That's pretty close to what I was hoping for/expecting. If I increase the tile overlaps for some of my smaller models, kind of like a small TTA in some ways, I see noticable improvement (about 0.01-0.02) so I suspect that proper TTA will be a larger boost.",
      "votes": null
    },
    {
      "id": "3070515",
      "postDate": "12/12/2024 19:02:57",
      "content": "<p>you should target +0.030 to 0.040 for ensemble+TTA+post-processing. That is quite a lot.</p>\n<p>2 models with 6xTTA + 1xOriginal  would use about 50 sec per tomography using 2xT4 GPU<br>\nThat is about 5hr 30min for 500 volumes</p>\n<p>This is without acceleration like torch compile or tensorRT which should further increase speed by x2.</p>",
      "rawMarkdown": "you should target +0.030 to 0.040 for ensemble+TTA+post-processing. That is quite a lot.\n\n2 models with 6xTTA + 1xOriginal  would use about 50 sec per tomography using 2xT4 GPU\nThat is about 5hr 30min for 500 volumes\n\nThis is without acceleration like torch compile or tensorRT which should further increase speed by x2.",
      "votes": null
    },
    {
      "id": "3070519",
      "postDate": "12/12/2024 19:22:52",
      "content": "<p>0.3 to 0.4 would be a huge boost given my current LB scores are already &gt; 0.6. maybe a 0.1 to 0.2 is more likely?</p>\n<p>I'm using Keras 3.0 w/JAX and I'm looking for ways to speed these models up. I have a lot to try here.</p>\n<p>Thanks for the advice on the 2 model ensemble. I am doing some validation on getting better learning rate schedules that help reduce over fitting for my ensemble models and will try this once I have this worked out :). </p>",
      "rawMarkdown": "0.3 to 0.4 would be a huge boost given my current LB scores are already > 0.6. maybe a 0.1 to 0.2 is more likely?\n\nI'm using Keras 3.0 w/JAX and I'm looking for ways to speed these models up. I have a lot to try here.\n\nThanks for the advice on the 2 model ensemble. I am doing some validation on getting better learning rate schedules that help reduce over fitting for my ensemble models and will try this once I have this worked out :).",
      "votes": null
    },
    {
      "id": "3073216",
      "postDate": "12/16/2024 07:30:52",
      "content": "<p>Why did my score drop after using TTA, and I only applied flipping?</p>",
      "rawMarkdown": "Why did my score drop after using TTA, and I only applied flipping?",
      "votes": null
    },
    {
      "id": "3073792",
      "postDate": "12/16/2024 23:47:48",
      "content": "<p><a href=\"https://www.kaggle.com/linheshen\" target=\"_blank\">@linheshen</a> I think this post might help you: <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/549744\" target=\"_blank\">flip discussion post</a>. I can't say for sure why you didn't get an improved score and I'm wrapping up some ensemble studies and starting to look at TTA, but I don't have results yet to speak of. Do you see this behavior in your local validation? Did you train with flips? I'd start with these questions and go from there.</p>",
      "rawMarkdown": "linheshen I think this post might help you: [flip discussion post](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/549744). I can't say for sure why you didn't get an improved score and I'm wrapping up some ensemble studies and starting to look at TTA, but I don't have results yet to speak of. Do you see this behavior in your local validation? Did you train with flips? I'd start with these questions and go from there.",
      "votes": null
    },
    {
      "id": "3074080",
      "postDate": "12/17/2024 09:53:08",
      "content": "<p>You are right, thank you very much for your suggestion. I am retraining the model now.</p>",
      "rawMarkdown": "You are right, thank you very much for your suggestion. I am retraining the model now.",
      "votes": null
    },
    {
      "id": "3074593",
      "postDate": "12/17/2024 20:50:20",
      "content": "<p>Good luck!</p>",
      "rawMarkdown": "Good luck!",
      "votes": null
    },
    {
      "id": "3074718",
      "postDate": "12/18/2024 01:29:19",
      "content": "<p>for TTA, one should work out the math.<br>\nlet num of hit, FP, and a1,b1 for TTA 1.<br>\nlet num of hit, FP, and a2,b2 for TTA 2.</p>\n<p>….</p>\n<p>if we combine TTA1 and 2,<br>\nthen combined num of hit, FP, = ….<br>\nthen recall, precision, fbeta = ….<br>\n(note: it may not be simple addition a prediction may occurs in both TTA and TTA2, in that case you have to use some probablistic modeling)</p>\n<p>there are some limits. e.g</p>\n<ul>\n<li>FP must be below xxx for TTA to work</li>\n<li>overlap in prediction must be at least xxx for TTA to work</li>\n</ul>\n<hr>\n<p>if you don't use math, then you can do experim,ents to collect the numbers. width sufficient statistics, you can predict the results on public and private score.</p>\n<hr>\n<p>nowadays, chatgpt can help</p>",
      "rawMarkdown": "for TTA, one should work out the math.\nlet num of hit, FP, and a1,b1 for TTA 1.\nlet num of hit, FP, and a2,b2 for TTA 2.\n\n....\n\nif we combine TTA1 and 2,\nthen combined num of hit, FP, = ....\nthen recall, precision, fbeta = ....\n(note: it may not be simple addition a prediction may occurs in both TTA and TTA2, in that case you have to use some probablistic modeling)\n\n\nthere are some limits. e.g\n- FP must be below xxx for TTA to work\n- overlap in prediction must be at least xxx for TTA to work\n\n---\n\nif you don't use math, then you can do experim,ents to collect the numbers. width sufficient statistics, you can predict the results on public and private score.\n\n---\n\nnowadays, chatgpt can help",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3067380,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "12/09/2024 08:01:09",
      "content": "<p>the first thing is to speed up your single model.<br>\ncheck the public code and post.<br>\nyou should be able to get about lb0.650 within 1.5 to 2 hours.</p>\n<p>distillation, etc. is only used as the last resort.</p>\n<p>if you cannot ensemble (and TTA) your score will not increase.<br>\nyou need these to get lb0.750 and beyond</p>",
      "votes": null,
      "replies": [
        {
          "id": 3067564,
          "author_name": "chemdatafarmer",
          "author_url": "",
          "post_date": "12/09/2024 12:11:57",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> ! I saw your pytorch code to speed up the connected components. I will spend some time to understand and incorporate that into my workflow. I think I can probably speed up how I generate the tiles too. This will give me more headroom for ensembling &amp; TTA. I'll try to find ways to make the model faster too.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3070229,
              "author_name": "chemdatafarmer",
              "author_url": "",
              "post_date": "12/12/2024 12:58:11",
              "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> you were right, ensembling can help the public score a lot! Finding the right thresholds helps a lot too. I still need to make the code faster so I can do TTA 😂. I've got a 7-fold ensemble running in ~11.5 hours haha…maybe 7 folds was a bit too much.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3070268,
                  "author_name": "andreizamfir",
                  "author_url": "",
                  "post_date": "12/12/2024 13:44:05",
                  "content": "<p>How much did ensembling improve your score?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3070333,
                      "author_name": "chemdatafarmer",
                      "author_url": "",
                      "post_date": "12/12/2024 15:13:07",
                      "content": "<p>Still doing some validation on that by checking the scores from the individual models with optimized thresholds, but it looks like it will be by about 0.05 or so.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3070339,
                          "author_name": "hengck23",
                          "author_url": "",
                          "post_date": "12/12/2024 15:17:44",
                          "content": "<p>'much' more than that is possible. it really depends on your model accuracy, tta, etc.</p>\n<p>if you wan to know the theoretical limit, google or ask chatgpt.<br>\ntrain accuracy is a weak upper bound</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 3070453,
                              "author_name": "chemdatafarmer",
                              "author_url": "",
                              "post_date": "12/12/2024 17:14:10",
                              "content": "<p>Thanks <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> ! I believe that's absolutely possible. My current models are likely a bit overfit. I'm not using much augmentation yet. This ensemble was also trained with a fixed learning rate schedule I optimized for one of the folds and I actually think on some of the folds, this caused quite a bit of overfitting. There's still a pretty large gap (about 0.1) from my CV F_Beta score and my LB F_Beta score. I will keep trying!</p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        },
                        {
                          "id": 3070340,
                          "author_name": "andreizamfir",
                          "author_url": "",
                          "post_date": "12/12/2024 15:17:48",
                          "content": "<p>Thanks, I haven't done ensembling yet as I ran out of GPU for the week, but with TTA I saw an increase of ~0.03.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 3070456,
                              "author_name": "chemdatafarmer",
                              "author_url": "",
                              "post_date": "12/12/2024 17:16:01",
                              "content": "<p><a href=\"https://www.kaggle.com/andreizamfir\" target=\"_blank\">@andreizamfir</a> good to know! That's pretty close to what I was hoping for/expecting. If I increase the tile overlaps for some of my smaller models, kind of like a small TTA in some ways, I see noticable improvement (about 0.01-0.02) so I suspect that proper TTA will be a larger boost.</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 3070515,
                                  "author_name": "hengck23",
                                  "author_url": "",
                                  "post_date": "12/12/2024 19:02:57",
                                  "content": "<p>you should target +0.030 to 0.040 for ensemble+TTA+post-processing. That is quite a lot.</p>\n<p>2 models with 6xTTA + 1xOriginal  would use about 50 sec per tomography using 2xT4 GPU<br>\nThat is about 5hr 30min for 500 volumes</p>\n<p>This is without acceleration like torch compile or tensorRT which should further increase speed by x2.</p>",
                                  "votes": null,
                                  "replies": [
                                    {
                                      "id": 3070519,
                                      "author_name": "chemdatafarmer",
                                      "author_url": "",
                                      "post_date": "12/12/2024 19:22:52",
                                      "content": "<p>0.3 to 0.4 would be a huge boost given my current LB scores are already &gt; 0.6. maybe a 0.1 to 0.2 is more likely?</p>\n<p>I'm using Keras 3.0 w/JAX and I'm looking for ways to speed these models up. I have a lot to try here.</p>\n<p>Thanks for the advice on the 2 model ensemble. I am doing some validation on getting better learning rate schedules that help reduce over fitting for my ensemble models and will try this once I have this worked out :). </p>",
                                      "votes": null,
                                      "replies": []
                                    }
                                  ]
                                }
                              ]
                            },
                            {
                              "id": 3073216,
                              "author_name": "linheshen",
                              "author_url": "",
                              "post_date": "12/16/2024 07:30:52",
                              "content": "<p>Why did my score drop after using TTA, and I only applied flipping?</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 3073792,
                                  "author_name": "chemdatafarmer",
                                  "author_url": "",
                                  "post_date": "12/16/2024 23:47:48",
                                  "content": "<p><a href=\"https://www.kaggle.com/linheshen\" target=\"_blank\">@linheshen</a> I think this post might help you: <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/549744\" target=\"_blank\">flip discussion post</a>. I can't say for sure why you didn't get an improved score and I'm wrapping up some ensemble studies and starting to look at TTA, but I don't have results yet to speak of. Do you see this behavior in your local validation? Did you train with flips? I'd start with these questions and go from there.</p>",
                                  "votes": null,
                                  "replies": [
                                    {
                                      "id": 3074080,
                                      "author_name": "linheshen",
                                      "author_url": "",
                                      "post_date": "12/17/2024 09:53:08",
                                      "content": "<p>You are right, thank you very much for your suggestion. I am retraining the model now.</p>",
                                      "votes": null,
                                      "replies": [
                                        {
                                          "id": 3074593,
                                          "author_name": "chemdatafarmer",
                                          "author_url": "",
                                          "post_date": "12/17/2024 20:50:20",
                                          "content": "<p>Good luck!</p>",
                                          "votes": null,
                                          "replies": [
                                            {
                                              "id": 3074718,
                                              "author_name": "hengck23",
                                              "author_url": "",
                                              "post_date": "12/18/2024 01:29:19",
                                              "content": "<p>for TTA, one should work out the math.<br>\nlet num of hit, FP, and a1,b1 for TTA 1.<br>\nlet num of hit, FP, and a2,b2 for TTA 2.</p>\n<p>….</p>\n<p>if we combine TTA1 and 2,<br>\nthen combined num of hit, FP, = ….<br>\nthen recall, precision, fbeta = ….<br>\n(note: it may not be simple addition a prediction may occurs in both TTA and TTA2, in that case you have to use some probablistic modeling)</p>\n<p>there are some limits. e.g</p>\n<ul>\n<li>FP must be below xxx for TTA to work</li>\n<li>overlap in prediction must be at least xxx for TTA to work</li>\n</ul>\n<hr>\n<p>if you don't use math, then you can do experim,ents to collect the numbers. width sufficient statistics, you can predict the results on public and private score.</p>\n<hr>\n<p>nowadays, chatgpt can help</p>",
                                              "votes": null,
                                              "replies": []
                                            }
                                          ]
                                        }
                                      ]
                                    }
                                  ]
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3067109": "Hey everyone,\n\nI'm currently working with single fold models and I'd like to find a way to ensemble my cross-validation models. My current train F_beta is 0.76, validation (two samples) is 0.75, and my LB score with this single fold model is 0.639. My hope is that building a model that's an ensemble of the cross-validation scores will help me reduce the gap between the leaderboard and my validation data. Note: it's entirely possible that this LB score is closer to my true cross-validation score and perhaps I'm validating on two 'easy' samples and I'm working on validating this.\n\nOf course I'm going to experiment with this on my own, but I'm looking to learn from the Kaggle community's experience here. When trying to keep your submissions < 12 hours, how do you usually approach this? I've done some reading and used Chat GPT to get me started. Here are 4 things I think sound reasonable and I'd like to hear your opinions/experiences.\n\n1) Average predictions from each of the folded models independently. This seems straightforward and I've used it before, but it could take some time. I've got some headroom (my current process takes about 6 hours to score), but I'd like to learn about other more efficient ways to do this.\n\n2) Weight interpolation. This seems interesting, performant, and straight-forward, but it's a bit of a new concept for me. Have you had success with this in the past? It seems logical, but it also seems like things could go wrong very quickly.\n\n3) Distillation. This seems like a neat way to get one smaller model that is a blend of the many, but I'm a little worried that it would be prone to overfitting since having a true train/test split here would be tricky with the number of samples we have. Since the models we're distilling from are from different folds, it's hard to generate a true holdout set here with only 7 samples.\n\n4) Use CV to identify the best learning rate schedule, then apply that schedule to all 7 samples. This sounds pretty easy, but best practices about how to validate this sort of thing (outside of trusting the leaderboard) seems kind of tough. There's a little more 'faith' behind this approach than I'd like to have.",
    "3067380": "the first thing is to speed up your single model.\ncheck the public code and post.\nyou should be able to get about lb0.650 within 1.5 to 2 hours.\n\ndistillation, etc. is only used as the last resort.\n\nif you cannot ensemble (and TTA) your score will not increase.\nyou need these to get lb0.750 and beyond",
    "3067564": "Thank you @hengck23 ! I saw your pytorch code to speed up the connected components. I will spend some time to understand and incorporate that into my workflow. I think I can probably speed up how I generate the tiles too. This will give me more headroom for ensembling & TTA. I'll try to find ways to make the model faster too.",
    "3070229": "hengck23 you were right, ensembling can help the public score a lot! Finding the right thresholds helps a lot too. I still need to make the code faster so I can do TTA 😂. I've got a 7-fold ensemble running in ~11.5 hours haha...maybe 7 folds was a bit too much.",
    "3070268": "How much did ensembling improve your score?",
    "3070333": "Still doing some validation on that by checking the scores from the individual models with optimized thresholds, but it looks like it will be by about 0.05 or so.",
    "3070339": "'much' more than that is possible. it really depends on your model accuracy, tta, etc.\n\nif you wan to know the theoretical limit, google or ask chatgpt.\ntrain accuracy is a weak upper bound",
    "3070340": "Thanks, I haven't done ensembling yet as I ran out of GPU for the week, but with TTA I saw an increase of ~0.03.",
    "3070453": "Thanks @hengck23 ! I believe that's absolutely possible. My current models are likely a bit overfit. I'm not using much augmentation yet. This ensemble was also trained with a fixed learning rate schedule I optimized for one of the folds and I actually think on some of the folds, this caused quite a bit of overfitting. There's still a pretty large gap (about 0.1) from my CV F_Beta score and my LB F_Beta score. I will keep trying!",
    "3070456": "andreizamfir good to know! That's pretty close to what I was hoping for/expecting. If I increase the tile overlaps for some of my smaller models, kind of like a small TTA in some ways, I see noticable improvement (about 0.01-0.02) so I suspect that proper TTA will be a larger boost.",
    "3070515": "you should target +0.030 to 0.040 for ensemble+TTA+post-processing. That is quite a lot.\n\n2 models with 6xTTA + 1xOriginal  would use about 50 sec per tomography using 2xT4 GPU\nThat is about 5hr 30min for 500 volumes\n\nThis is without acceleration like torch compile or tensorRT which should further increase speed by x2.",
    "3070519": "0.3 to 0.4 would be a huge boost given my current LB scores are already > 0.6. maybe a 0.1 to 0.2 is more likely?\n\nI'm using Keras 3.0 w/JAX and I'm looking for ways to speed these models up. I have a lot to try here.\n\nThanks for the advice on the 2 model ensemble. I am doing some validation on getting better learning rate schedules that help reduce over fitting for my ensemble models and will try this once I have this worked out :).",
    "3073216": "Why did my score drop after using TTA, and I only applied flipping?",
    "3073792": "linheshen I think this post might help you: [flip discussion post](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/549744). I can't say for sure why you didn't get an improved score and I'm wrapping up some ensemble studies and starting to look at TTA, but I don't have results yet to speak of. Do you see this behavior in your local validation? Did you train with flips? I'd start with these questions and go from there.",
    "3074080": "You are right, thank you very much for your suggestion. I am retraining the model now.",
    "3074593": "Good luck!",
    "3074718": "for TTA, one should work out the math.\nlet num of hit, FP, and a1,b1 for TTA 1.\nlet num of hit, FP, and a2,b2 for TTA 2.\n\n....\n\nif we combine TTA1 and 2,\nthen combined num of hit, FP, = ....\nthen recall, precision, fbeta = ....\n(note: it may not be simple addition a prediction may occurs in both TTA and TTA2, in that case you have to use some probablistic modeling)\n\n\nthere are some limits. e.g\n- FP must be below xxx for TTA to work\n- overlap in prediction must be at least xxx for TTA to work\n\n---\n\nif you don't use math, then you can do experim,ents to collect the numbers. width sufficient statistics, you can predict the results on public and private score.\n\n---\n\nnowadays, chatgpt can help"
  },
  "source": "meta"
}