{
  "id": 555670,
  "title": "Single Model vs. Ensemble/TTA Output",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/555670",
  "author_name": "David List",
  "post_date": "2025-01-08T17:10:36.043000",
  "votes": 13,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Thought I would share.  These are images of the raw and softmax thyroglobulin output for single and ensemble/tta models of TS_69_2.  It's a bit subtle, but in the softmax at least, the ensemble/tta predictions are a bit cleaner than the single model predictions:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10704200%2Fd0421b55ae4e82327ed61119c8374c45%2FSingleVsEnsemble.png?generation=1736355778043945&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 3091684,
      "postDate": "2025-01-08T17:10:36.043Z",
      "content": "<p>Thought I would share.  These are images of the raw and softmax thyroglobulin output for single and ensemble/tta models of TS_69_2.  It's a bit subtle, but in the softmax at least, the ensemble/tta predictions are a bit cleaner than the single model predictions:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10704200%2Fd0421b55ae4e82327ed61119c8374c45%2FSingleVsEnsemble.png?generation=1736355778043945&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Thought I would share.  These are images of the raw and softmax thyroglobulin output for single and ensemble/tta models of TS_69_2.  It's a bit subtle, but in the softmax at least, the ensemble/tta predictions are a bit cleaner than the single model predictions:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10704200%2Fd0421b55ae4e82327ed61119c8374c45%2FSingleVsEnsemble.png?generation=1736355778043945&alt=media)",
      "votes": 13
    },
    {
      "id": 3091816,
      "postDate": "2025-01-08T19:35:57.443Z",
      "content": "<p>I agree.  Those images are alot cleaner using Ensemble/TTA softmax. Thank you for sharing.</p>",
      "rawMarkdown": "I agree.  Those images are alot cleaner using Ensemble/TTA softmax. Thank you for sharing.",
      "votes": 1,
      "replies": [
        {
          "id": 3107981,
          "postDate": "2025-01-27T07:52:49.140Z",
          "content": "<p><a href=\"https://www.kaggle.com/johnkennedymlops\" target=\"_blank\">@johnkennedymlops</a> i know that is a dummy question,  Ensemble/TTA softmax is that means some kind of augmentation or ensembling models , i'm confused</p>",
          "rawMarkdown": "@johnkennedymlops i know that is a dummy question,  Ensemble/TTA softmax is that means some kind of augmentation or ensembling models , i'm confused",
          "replies": [
            {
              "id": 3107996,
              "postDate": "2025-01-27T08:32:27.097Z",
              "content": "<p>Ensemble/TTA Softmax means 1) It's an ensemble of multiple models.  2) TTA i(Test Time Augmentation) has been applied. (Look this up if you haven't heard of it) 3) Softmax, which scales the outputs to relative probabilities, has been applied</p>",
              "rawMarkdown": "Ensemble/TTA Softmax means 1) It's an ensemble of multiple models.  2) TTA i(Test Time Augmentation) has been applied. (Look this up if you haven't heard of it) 3) Softmax, which scales the outputs to relative probabilities, has been applied",
              "votes": 1
            },
            {
              "id": 3108110,
              "postDate": "2025-01-27T11:44:47.953Z",
              "content": "<p>i get it now , you're ensembling many models and applying augmentation during inference( not during training ) and applied a softmax at the predictions to choose the predictions with highest probabilities</p>",
              "rawMarkdown": "i get it now , you're ensembling many models and applying augmentation during inference( not during training ) and applied a softmax at the predictions to choose the predictions with highest probabilities"
            }
          ]
        }
      ]
    },
    {
      "id": 3091783,
      "postDate": "2025-01-08T18:43:28.127Z",
      "content": "<p>This is neat, thanks for sharing! I'm still only seeing marginal benefits from ensembling and TTA. My best ensemble w/TTA scores at 0.742 and my best single model with no TTA scores at 0.736. Validation metrics for them are more or less comparable.</p>",
      "rawMarkdown": "This is neat, thanks for sharing! I'm still only seeing marginal benefits from ensembling and TTA. My best ensemble w/TTA scores at 0.742 and my best single model with no TTA scores at 0.736. Validation metrics for them are more or less comparable.",
      "votes": 1,
      "replies": [
        {
          "id": 3091792,
          "postDate": "2025-01-08T18:56:33.690Z",
          "content": "<p>Interesting.  I went from 0.704 to 0.734 so it was pretty huge for me.  Could be the ensemble was \"fixing\" some other issues my model had that yours didn't.</p>",
          "rawMarkdown": "Interesting.  I went from 0.704 to 0.734 so it was pretty huge for me.  Could be the ensemble was \"fixing\" some other issues my model had that yours didn't.",
          "votes": 1,
          "replies": [
            {
              "id": 3091804,
              "postDate": "2025-01-08T19:14:45.477Z",
              "content": "<p>I think it's also possible my models are a little biased to false positives (which I'm working to see if I can fix) and doing TTA fixes some things, but also adds more false positives. If I get my false positives down a bit, I wonder if I'll see a jump with TTA? I'm working on that now.</p>",
              "rawMarkdown": "I think it's also possible my models are a little biased to false positives (which I'm working to see if I can fix) and doing TTA fixes some things, but also adds more false positives. If I get my false positives down a bit, I wonder if I'll see a jump with TTA? I'm working on that now.",
              "votes": 1
            },
            {
              "id": 3091815,
              "postDate": "2025-01-08T19:35:51.967Z",
              "content": "<p>When it comes to ensembling, I believe variety is the driving factor, you need solutions that complement eachother. TTA can also be aggregated in multiple ways and I believe it’s something to look into and experiment with.</p>",
              "rawMarkdown": "When it comes to ensembling, I believe variety is the driving factor, you need solutions that complement eachother. TTA can also be aggregated in multiple ways and I believe it’s something to look into and experiment with.",
              "votes": 2
            },
            {
              "id": 3091840,
              "postDate": "2025-01-08T20:30:14.260Z",
              "content": "<p>You're right, perhaps I need to rethink my ensembling strategy :). I've mostly been using CV folds for the ensembles, but it's possible they're just reinforcing their own strengths and biases.</p>\n<p>I've done pretty basic TTA work up until now, there's probably more targeted TTAs I can try explicitly designed to help with my model's current weaknesses.</p>",
              "rawMarkdown": "You're right, perhaps I need to rethink my ensembling strategy :). I've mostly been using CV folds for the ensembles, but it's possible they're just reinforcing their own strengths and biases.\n\nI've done pretty basic TTA work up until now, there's probably more targeted TTAs I can try explicitly designed to help with my model's current weaknesses."
            },
            {
              "id": 3092667,
              "postDate": "2025-01-09T20:23:43.803Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 3092668,
              "postDate": "2025-01-09T20:24:23.623Z",
              "content": "<p><a href=\"https://www.kaggle.com/chemdatafarmer\" target=\"_blank\">@chemdatafarmer</a>, I agree. Walking the precision/recall curve to get the right balance of fp's is way more important than TTA.</p>",
              "rawMarkdown": "@chemdatafarmer, I agree. Walking the precision/recall curve to get the right balance of fp's is way more important than TTA.",
              "votes": 3
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3091816,
      "author_name": "JohnKennedyMLOps",
      "author_url": "",
      "post_date": "2025-01-08T19:35:57.443000",
      "content": "<p>I agree.  Those images are alot cleaner using Ensemble/TTA softmax. Thank you for sharing.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3107981,
          "author_name": "work work",
          "author_url": "",
          "post_date": "2025-01-27T07:52:49.140000",
          "content": "<p><a href=\"https://www.kaggle.com/johnkennedymlops\" target=\"_blank\">@johnkennedymlops</a> i know that is a dummy question,  Ensemble/TTA softmax is that means some kind of augmentation or ensembling models , i'm confused</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3107996,
              "author_name": "David List",
              "author_url": "",
              "post_date": "2025-01-27T08:32:27.097000",
              "content": "<p>Ensemble/TTA Softmax means 1) It's an ensemble of multiple models.  2) TTA i(Test Time Augmentation) has been applied. (Look this up if you haven't heard of it) 3) Softmax, which scales the outputs to relative probabilities, has been applied</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3108110,
              "author_name": "work work",
              "author_url": "",
              "post_date": "2025-01-27T11:44:47.953000",
              "content": "<p>i get it now , you're ensembling many models and applying augmentation during inference( not during training ) and applied a softmax at the predictions to choose the predictions with highest probabilities</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3091783,
      "author_name": "chemdatafarmer",
      "author_url": "",
      "post_date": "2025-01-08T18:43:28.127000",
      "content": "<p>This is neat, thanks for sharing! I'm still only seeing marginal benefits from ensembling and TTA. My best ensemble w/TTA scores at 0.742 and my best single model with no TTA scores at 0.736. Validation metrics for them are more or less comparable.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3091792,
          "author_name": "David List",
          "author_url": "",
          "post_date": "2025-01-08T18:56:33.690000",
          "content": "<p>Interesting.  I went from 0.704 to 0.734 so it was pretty huge for me.  Could be the ensemble was \"fixing\" some other issues my model had that yours didn't.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3091804,
              "author_name": "chemdatafarmer",
              "author_url": "",
              "post_date": "2025-01-08T19:14:45.477000",
              "content": "<p>I think it's also possible my models are a little biased to false positives (which I'm working to see if I can fix) and doing TTA fixes some things, but also adds more false positives. If I get my false positives down a bit, I wonder if I'll see a jump with TTA? I'm working on that now.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3091815,
              "author_name": "Andrei Zamfir",
              "author_url": "",
              "post_date": "2025-01-08T19:35:51.967000",
              "content": "<p>When it comes to ensembling, I believe variety is the driving factor, you need solutions that complement eachother. TTA can also be aggregated in multiple ways and I believe it’s something to look into and experiment with.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3091840,
              "author_name": "chemdatafarmer",
              "author_url": "",
              "post_date": "2025-01-08T20:30:14.260000",
              "content": "<p>You're right, perhaps I need to rethink my ensembling strategy :). I've mostly been using CV folds for the ensembles, but it's possible they're just reinforcing their own strengths and biases.</p>\n<p>I've done pretty basic TTA work up until now, there's probably more targeted TTAs I can try explicitly designed to help with my model's current weaknesses.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3092667,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-01-09T20:23:43.803000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3092668,
              "author_name": "David List",
              "author_url": "",
              "post_date": "2025-01-09T20:24:23.623000",
              "content": "<p><a href=\"https://www.kaggle.com/chemdatafarmer\" target=\"_blank\">@chemdatafarmer</a>, I agree. Walking the precision/recall curve to get the right balance of fp's is way more important than TTA.</p>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3091684": "Thought I would share.  These are images of the raw and softmax thyroglobulin output for single and ensemble/tta models of TS_69_2.  It's a bit subtle, but in the softmax at least, the ensemble/tta predictions are a bit cleaner than the single model predictions:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10704200%2Fd0421b55ae4e82327ed61119c8374c45%2FSingleVsEnsemble.png?generation=1736355778043945&alt=media)",
    "3091816": "I agree.  Those images are alot cleaner using Ensemble/TTA softmax. Thank you for sharing.",
    "3091783": "This is neat, thanks for sharing! I'm still only seeing marginal benefits from ensembling and TTA. My best ensemble w/TTA scores at 0.742 and my best single model with no TTA scores at 0.736. Validation metrics for them are more or less comparable."
  }
}