{
  "id": 133318,
  "title": "The right way to ensemble models",
  "url": "/competitions/deepfake-detection-challenge/discussion/133318",
  "author_name": "",
  "post_date": "2020-03-02T03:04:01.884767300Z",
  "votes": null,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Can someone suggest how to ensemble models in the right way? I can not get any improvement by averaging</p>",
  "messages": [
    {
      "id": "761012",
      "postDate": "03/02/2020 03:04:01",
      "content": "<p>Can someone suggest how to ensemble models in the right way? I can not get any improvement by averaging</p>",
      "rawMarkdown": "Can someone suggest how to ensemble models in the right way? I can not get any improvement by averaging",
      "votes": null
    },
    {
      "id": "761021",
      "postDate": "03/02/2020 03:24:56",
      "content": "<p>It would help if you describe how you do it now... Its better to use models of different architecture, resolution and augmentation. Averaging usually works very well if your models are good and not over fitting. There are much more complicated ways to do it but I am not sure they will help if averaging doesn't. </p>",
      "rawMarkdown": "It would help if you describe how you do it now... Its better to use models of different architecture, resolution and augmentation. Averaging usually works very well if your models are good and not over fitting. There are much more complicated ways to do it but I am not sure they will help if averaging doesn't.",
      "votes": null
    },
    {
      "id": "761041",
      "postDate": "03/02/2020 04:12:32",
      "content": "<p>Thanks! I`m using models of different architecture, but image resolution and augmentations are the same. I tried to average only to models because I have only to different architectures, one scored 0.33 other 0.329, averaging gives 0.325 (Maybe it is good). Both models are trained on the same fold because I suspect other splits lead to overfitting </p>",
      "rawMarkdown": "Thanks! I`m using models of different architecture, but image resolution and augmentations are the same. I tried to average only to models because I have only to different architectures, one scored 0.33 other 0.329, averaging gives 0.325 (Maybe it is good). Both models are trained on the same fold because I suspect other splits lead to overfitting",
      "votes": null
    },
    {
      "id": "761060",
      "postDate": "03/02/2020 04:49:49",
      "content": "<p>That is not bad. Ensembling is more trial and error than anything else. Sometimes ensembling two runs with different seeds also yields nice results but what you did is correct. In non code competitions people ensamble over 10 models but here it won't work... </p>",
      "rawMarkdown": "That is not bad. Ensembling is more trial and error than anything else. Sometimes ensembling two runs with different seeds also yields nice results but what you did is correct. In non code competitions people ensamble over 10 models but here it won't work...",
      "votes": null
    },
    {
      "id": "761077",
      "postDate": "03/02/2020 05:16:32",
      "content": "<p>OK, Thanks!</p>",
      "rawMarkdown": "OK, Thanks!",
      "votes": null
    },
    {
      "id": "761310",
      "postDate": "03/02/2020 11:31:01",
      "content": "<p>A simple improvement is to do weighted average, giving the better model more weight </p>",
      "rawMarkdown": "A simple improvement is to do weighted average, giving the better model more weight",
      "votes": null
    },
    {
      "id": "761684",
      "postDate": "03/02/2020 21:04:33",
      "content": "<p>Try weighted average by discretized probability score.</p>",
      "rawMarkdown": "Try weighted average by discretized probability score.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 761021,
      "author_name": "moshel",
      "author_url": "",
      "post_date": "03/02/2020 03:24:56",
      "content": "<p>It would help if you describe how you do it now... Its better to use models of different architecture, resolution and augmentation. Averaging usually works very well if your models are good and not over fitting. There are much more complicated ways to do it but I am not sure they will help if averaging doesn't. </p>",
      "votes": null,
      "replies": [
        {
          "id": 761041,
          "author_name": "azamatk",
          "author_url": "",
          "post_date": "03/02/2020 04:12:32",
          "content": "<p>Thanks! I`m using models of different architecture, but image resolution and augmentations are the same. I tried to average only to models because I have only to different architectures, one scored 0.33 other 0.329, averaging gives 0.325 (Maybe it is good). Both models are trained on the same fold because I suspect other splits lead to overfitting </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 761060,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "03/02/2020 04:49:49",
          "content": "<p>That is not bad. Ensembling is more trial and error than anything else. Sometimes ensembling two runs with different seeds also yields nice results but what you did is correct. In non code competitions people ensamble over 10 models but here it won't work... </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 761077,
          "author_name": "azamatk",
          "author_url": "",
          "post_date": "03/02/2020 05:16:32",
          "content": "<p>OK, Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 761310,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "03/02/2020 11:31:01",
          "content": "<p>A simple improvement is to do weighted average, giving the better model more weight </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 761684,
          "author_name": "maralski",
          "author_url": "",
          "post_date": "03/02/2020 21:04:33",
          "content": "<p>Try weighted average by discretized probability score.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "761012": "Can someone suggest how to ensemble models in the right way? I can not get any improvement by averaging",
    "761021": "It would help if you describe how you do it now... Its better to use models of different architecture, resolution and augmentation. Averaging usually works very well if your models are good and not over fitting. There are much more complicated ways to do it but I am not sure they will help if averaging doesn't.",
    "761041": "Thanks! I`m using models of different architecture, but image resolution and augmentations are the same. I tried to average only to models because I have only to different architectures, one scored 0.33 other 0.329, averaging gives 0.325 (Maybe it is good). Both models are trained on the same fold because I suspect other splits lead to overfitting",
    "761060": "That is not bad. Ensembling is more trial and error than anything else. Sometimes ensembling two runs with different seeds also yields nice results but what you did is correct. In non code competitions people ensamble over 10 models but here it won't work...",
    "761077": "OK, Thanks!",
    "761310": "A simple improvement is to do weighted average, giving the better model more weight",
    "761684": "Try weighted average by discretized probability score."
  },
  "source": "meta"
}