{
  "id": 146024,
  "title": "Did anyone get medal with single-model?",
  "url": "/competitions/deepfake-detection-challenge/discussion/146024",
  "author_name": "",
  "post_date": "2020-04-25T14:40:40.937374600Z",
  "votes": 6,
  "comment_count": 13,
  "views": 0,
  "content": "<p>I incoporated ensemble models for private test as I thought it alleviate overfitting.</p>\n\n<p>And I am curious about any person get into medal zone with not ensemble but single model.</p>\n\n<p>So then could you share your model and your generalization strategy?</p>",
  "messages": [
    {
      "id": "820567",
      "postDate": "04/25/2020 14:40:40",
      "content": "<p>I incoporated ensemble models for private test as I thought it alleviate overfitting.</p>\n\n<p>And I am curious about any person get into medal zone with not ensemble but single model.</p>\n\n<p>So then could you share your model and your generalization strategy?</p>",
      "rawMarkdown": "I incoporated ensemble models for private test as I thought it alleviate overfitting.\n\nAnd I am curious about any person get into medal zone with not ensemble but single model.\n\nSo then could you share your model and your generalization strategy?",
      "votes": null
    },
    {
      "id": "821132",
      "postDate": "04/25/2020 23:58:03",
      "content": "<p>You could call the 43rd place solution a single model. They did ensemble multiple checkpoints though. They used EfficientNetB4 with added attention layers and wrote a full paper on it.</p>\n\n<p>Discussion Post:\n<a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/145841\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/145841</a></p>\n\n<p>Paper:\n<a href=\"https://arxiv.org/pdf/2004.07676.pdf\">https://arxiv.org/pdf/2004.07676.pdf</a></p>\n\n<p>Notebook:\n<a href=\"https://www.kaggle.com/nicobonne/43-rank-ispl-ensamble-10-model\">https://www.kaggle.com/nicobonne/43-rank-ispl-ensamble-10-model</a></p>\n\n<p>Github repo:\n<a href=\"https://github.com/polimi-ispl/icpr2020dfdc\">https://github.com/polimi-ispl/icpr2020dfdc</a></p>\n\n<p>Hope this helps!</p>",
      "rawMarkdown": "You could call the 43rd place solution a single model. They did ensemble multiple checkpoints though. They used EfficientNetB4 with added attention layers and wrote a full paper on it.\n\nDiscussion Post:\nhttps://www.kaggle.com/c/deepfake-detection-challenge/discussion/145841\n\nPaper:\nhttps://arxiv.org/pdf/2004.07676.pdf\n\nNotebook:\nhttps://www.kaggle.com/nicobonne/43-rank-ispl-ensamble-10-model\n\nGithub repo:\nhttps://github.com/polimi-ispl/icpr2020dfdc\n\nHope this helps!",
      "votes": null
    },
    {
      "id": "821816",
      "postDate": "04/26/2020 12:46:12",
      "content": "<p>Thank you for you reply!</p>\n\n<p>Efficientnet is very remarkable for this competition besides widely known other-task.\nIt's very interesting they use just EfficientnetB4 for back-bone though train variants( attention / siamese) for ensembel.</p>\n\n<p>I guess they could have achieved better perfomance if they used more model(b5/b6...) for their ensemble.</p>",
      "rawMarkdown": "Thank you for you reply!\n\nEfficientnet is very remarkable for this competition besides widely known other-task.\nIt's very interesting they use just EfficientnetB4 for back-bone though train variants( attention / siamese) for ensembel.\n\nI guess they could have achieved better perfomance if they used more model(b5/b6...) for their ensemble.",
      "votes": null
    },
    {
      "id": "823713",
      "postDate": "04/27/2020 20:31:25",
      "content": "<p>Literally in the last week of the competition, I tried a generalization technique which significantly improved my score. I was working on it through the week nights, and as such I ran out of time to generate proper ensemble solutions :(\nBut, since it was better than the others, I ended up making one of my two submissions a single model which scored 0.30848 in public LB and 0.48569 in private LB (it would place 47th place in private LB).\nThe only other submission I managed to train and submit within the deadline, was an ensemble of the first submission model, with another training run of the same model with a different random seed. The ensemble public LB was 0.30313 (66th place) and the private LB was 0.47007 ( 21st place ).\nI have to say that given the quick improvement in score for the simple two model ensemble, I'm quite annoyed that I didn't get a chance to prepare a proper ensemble.</p>",
      "rawMarkdown": "Literally in the last week of the competition, I tried a generalization technique which significantly improved my score. I was working on it through the week nights, and as such I ran out of time to generate proper ensemble solutions :(\nBut, since it was better than the others, I ended up making one of my two submissions a single model which scored 0.30848 in public LB and 0.48569 in private LB (it would place 47th place in private LB).\nThe only other submission I managed to train and submit within the deadline, was an ensemble of the first submission model, with another training run of the same model with a different random seed. The ensemble public LB was 0.30313 (66th place) and the private LB was 0.47007 ( 21st place ).\nI have to say that given the quick improvement in score for the simple two model ensemble, I'm quite annoyed that I didn't get a chance to prepare a proper ensemble.",
      "votes": null
    },
    {
      "id": "823884",
      "postDate": "04/28/2020 01:34:23",
      "content": "<p>Congrats to your decent outcome though you didn't have a chances to test your ensemble!</p>\n\n<p>Ensemble was so effective also for this competition especially enhancing generalization power \nconsidering all experience shraed on disscussion board.</p>\n\n<p>So, what kind of model did you use for your backbone? I'm curious about that.</p>",
      "rawMarkdown": "Congrats to your decent outcome though you didn't have a chances to test your ensemble!\n\nEnsemble was so effective also for this competition especially enhancing generalization power \nconsidering all experience shraed on disscussion board.\n\nSo, what kind of model did you use for your backbone? I'm curious about that.",
      "votes": null
    },
    {
      "id": "823908",
      "postDate": "04/28/2020 02:14:08",
      "content": "<p>Thanks. I used EfficientNet B7 with input size 224</p>",
      "rawMarkdown": "Thanks. I used EfficientNet B7 with input size 224",
      "votes": null
    },
    {
      "id": "824108",
      "postDate": "04/28/2020 06:50:06",
      "content": "<p>Our team decided to go for just a single model. Maybe if we had an ensemble, we would have reached a higher place due to generalization. Single model is a effnet-b4\nSome tweaks: \n- Range of (heavy) augmentations used\n- A less risky submission with probabilities clipped between 0.15 and 0.9\n- Clever face detection and multiple actor tracking system</p>",
      "rawMarkdown": "Our team decided to go for just a single model. Maybe if we had an ensemble, we would have reached a higher place due to generalization. Single model is a effnet-b4\nSome tweaks: \n- Range of (heavy) augmentations used\n- A less risky submission with probabilities clipped between 0.15 and 0.9\n- Clever face detection and multiple actor tracking system",
      "votes": null
    },
    {
      "id": "824274",
      "postDate": "04/28/2020 09:10:49",
      "content": "<p>wow...! </p>",
      "rawMarkdown": "wow...!",
      "votes": null
    },
    {
      "id": "824275",
      "postDate": "04/28/2020 09:13:10",
      "content": "<p>Stunning result... You get generalized model without ensemble! Would you mind if you share the technique?!</p>",
      "rawMarkdown": "Stunning result... You get generalized model without ensemble! Would you mind if you share the technique?!",
      "votes": null
    },
    {
      "id": "824280",
      "postDate": "04/28/2020 09:15:04",
      "content": "<p>Default Input size for b7 is 600. \nDo you have any special reason to use input size 224?</p>\n\n<p>it's interesting \nbecause I thought larger input-size have benefit to detect blending boundary.</p>",
      "rawMarkdown": "Default Input size for b7 is 600. \nDo you have any special reason to use input size 224?\n\nit's interesting \nbecause I thought larger input-size have benefit to detect blending boundary.",
      "votes": null
    },
    {
      "id": "824283",
      "postDate": "04/28/2020 09:20:35",
      "content": "<p>Effnet worked superbly for this competition.</p>\n\n<p>It's very nice to achiev high-rank with single model.</p>\n\n<p>Would you share multiple face tracking system in detail?</p>",
      "rawMarkdown": "Effnet worked superbly for this competition.\n\nIt's very nice to achiev high-rank with single model.\n\nWould you share multiple face tracking system in detail?",
      "votes": null
    },
    {
      "id": "825224",
      "postDate": "04/28/2020 21:29:26",
      "content": "<p>Thanks! True that, effnets deemed to be more superior for us more than any other model. To the point I was speculating most teams were using them too.</p>\n\n<p>Regarding the multiple face tracking system - credits goes to my team-mate <a href=\"/yannael\">@yannael</a>, so forgive me if there is some faults in my somewhat complicated description.</p>\n\n<p>Briefly, we track a person based on the bounding boxes of the faces. Our goal is to give every person a unique ID based on these boxes. We define a threshold distance of around ~100 pixels between the box coordinates. If that threshold is surpassed, a new ID box is created assuming there is a second person in the video. Then for each consecutive frame, we associate the new bounding box to their IDs based on the minimal distance to the last boxes from the previous frame.</p>",
      "rawMarkdown": "Thanks! True that, effnets deemed to be more superior for us more than any other model. To the point I was speculating most teams were using them too.\n\nRegarding the multiple face tracking system - credits goes to my team-mate @yannael, so forgive me if there is some faults in my somewhat complicated description.\n\nBriefly, we track a person based on the bounding boxes of the faces. Our goal is to give every person a unique ID based on these boxes. We define a threshold distance of around ~100 pixels between the box coordinates. If that threshold is surpassed, a new ID box is created assuming there is a second person in the video. Then for each consecutive frame, we associate the new bounding box to their IDs based on the minimal distance to the last boxes from the previous frame.",
      "votes": null
    },
    {
      "id": "825484",
      "postDate": "04/29/2020 03:25:40",
      "content": "<p>The original face sizes vary. I was concerned about up-scaling smallish faces too much and having them look fake due to upscaling artifacts, etc. I was trying to pick a size to minimize the average amount of up-scaling required. I tried some smaller sizes and some larger sizes and this size worked well.</p>",
      "rawMarkdown": "The original face sizes vary. I was concerned about up-scaling smallish faces too much and having them look fake due to upscaling artifacts, etc. I was trying to pick a size to minimize the average amount of up-scaling required. I tried some smaller sizes and some larger sizes and this size worked well.",
      "votes": null
    },
    {
      "id": "829790",
      "postDate": "05/02/2020 05:22:49",
      "content": "<p>On the other hand I concerned about resizing to smaller size compared to original might destroy manipulation artifacts. I tested input size 256-640 and conclude to use 512 resolution.\nNow I read your idea and there might be pros and cons of either resizing larger or smaller.</p>",
      "rawMarkdown": "On the other hand I concerned about resizing to smaller size compared to original might destroy manipulation artifacts. I tested input size 256-640 and conclude to use 512 resolution.\nNow I read your idea and there might be pros and cons of either resizing larger or smaller.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 821132,
      "author_name": "carlolepelaars",
      "author_url": "",
      "post_date": "04/25/2020 23:58:03",
      "content": "<p>You could call the 43rd place solution a single model. They did ensemble multiple checkpoints though. They used EfficientNetB4 with added attention layers and wrote a full paper on it.</p>\n\n<p>Discussion Post:\n<a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/145841\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/145841</a></p>\n\n<p>Paper:\n<a href=\"https://arxiv.org/pdf/2004.07676.pdf\">https://arxiv.org/pdf/2004.07676.pdf</a></p>\n\n<p>Notebook:\n<a href=\"https://www.kaggle.com/nicobonne/43-rank-ispl-ensamble-10-model\">https://www.kaggle.com/nicobonne/43-rank-ispl-ensamble-10-model</a></p>\n\n<p>Github repo:\n<a href=\"https://github.com/polimi-ispl/icpr2020dfdc\">https://github.com/polimi-ispl/icpr2020dfdc</a></p>\n\n<p>Hope this helps!</p>",
      "votes": null,
      "replies": [
        {
          "id": 821816,
          "author_name": "gwsong",
          "author_url": "",
          "post_date": "04/26/2020 12:46:12",
          "content": "<p>Thank you for you reply!</p>\n\n<p>Efficientnet is very remarkable for this competition besides widely known other-task.\nIt's very interesting they use just EfficientnetB4 for back-bone though train variants( attention / siamese) for ensembel.</p>\n\n<p>I guess they could have achieved better perfomance if they used more model(b5/b6...) for their ensemble.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 823713,
      "author_name": "cesb45",
      "author_url": "",
      "post_date": "04/27/2020 20:31:25",
      "content": "<p>Literally in the last week of the competition, I tried a generalization technique which significantly improved my score. I was working on it through the week nights, and as such I ran out of time to generate proper ensemble solutions :(\nBut, since it was better than the others, I ended up making one of my two submissions a single model which scored 0.30848 in public LB and 0.48569 in private LB (it would place 47th place in private LB).\nThe only other submission I managed to train and submit within the deadline, was an ensemble of the first submission model, with another training run of the same model with a different random seed. The ensemble public LB was 0.30313 (66th place) and the private LB was 0.47007 ( 21st place ).\nI have to say that given the quick improvement in score for the simple two model ensemble, I'm quite annoyed that I didn't get a chance to prepare a proper ensemble.</p>",
      "votes": null,
      "replies": [
        {
          "id": 823884,
          "author_name": "gwsong",
          "author_url": "",
          "post_date": "04/28/2020 01:34:23",
          "content": "<p>Congrats to your decent outcome though you didn't have a chances to test your ensemble!</p>\n\n<p>Ensemble was so effective also for this competition especially enhancing generalization power \nconsidering all experience shraed on disscussion board.</p>\n\n<p>So, what kind of model did you use for your backbone? I'm curious about that.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 823908,
          "author_name": "cesb45",
          "author_url": "",
          "post_date": "04/28/2020 02:14:08",
          "content": "<p>Thanks. I used EfficientNet B7 with input size 224</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 824275,
          "author_name": "seriousran",
          "author_url": "",
          "post_date": "04/28/2020 09:13:10",
          "content": "<p>Stunning result... You get generalized model without ensemble! Would you mind if you share the technique?!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 824280,
          "author_name": "gwsong",
          "author_url": "",
          "post_date": "04/28/2020 09:15:04",
          "content": "<p>Default Input size for b7 is 600. \nDo you have any special reason to use input size 224?</p>\n\n<p>it's interesting \nbecause I thought larger input-size have benefit to detect blending boundary.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 825484,
          "author_name": "cesb45",
          "author_url": "",
          "post_date": "04/29/2020 03:25:40",
          "content": "<p>The original face sizes vary. I was concerned about up-scaling smallish faces too much and having them look fake due to upscaling artifacts, etc. I was trying to pick a size to minimize the average amount of up-scaling required. I tried some smaller sizes and some larger sizes and this size worked well.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 829790,
          "author_name": "gwsong",
          "author_url": "",
          "post_date": "05/02/2020 05:22:49",
          "content": "<p>On the other hand I concerned about resizing to smaller size compared to original might destroy manipulation artifacts. I tested input size 256-640 and conclude to use 512 resolution.\nNow I read your idea and there might be pros and cons of either resizing larger or smaller.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 824108,
      "author_name": "rafiko1",
      "author_url": "",
      "post_date": "04/28/2020 06:50:06",
      "content": "<p>Our team decided to go for just a single model. Maybe if we had an ensemble, we would have reached a higher place due to generalization. Single model is a effnet-b4\nSome tweaks: \n- Range of (heavy) augmentations used\n- A less risky submission with probabilities clipped between 0.15 and 0.9\n- Clever face detection and multiple actor tracking system</p>",
      "votes": null,
      "replies": [
        {
          "id": 824274,
          "author_name": "seriousran",
          "author_url": "",
          "post_date": "04/28/2020 09:10:49",
          "content": "<p>wow...! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 824283,
          "author_name": "gwsong",
          "author_url": "",
          "post_date": "04/28/2020 09:20:35",
          "content": "<p>Effnet worked superbly for this competition.</p>\n\n<p>It's very nice to achiev high-rank with single model.</p>\n\n<p>Would you share multiple face tracking system in detail?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 825224,
          "author_name": "rafiko1",
          "author_url": "",
          "post_date": "04/28/2020 21:29:26",
          "content": "<p>Thanks! True that, effnets deemed to be more superior for us more than any other model. To the point I was speculating most teams were using them too.</p>\n\n<p>Regarding the multiple face tracking system - credits goes to my team-mate <a href=\"/yannael\">@yannael</a>, so forgive me if there is some faults in my somewhat complicated description.</p>\n\n<p>Briefly, we track a person based on the bounding boxes of the faces. Our goal is to give every person a unique ID based on these boxes. We define a threshold distance of around ~100 pixels between the box coordinates. If that threshold is surpassed, a new ID box is created assuming there is a second person in the video. Then for each consecutive frame, we associate the new bounding box to their IDs based on the minimal distance to the last boxes from the previous frame.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "820567": "I incoporated ensemble models for private test as I thought it alleviate overfitting.\n\nAnd I am curious about any person get into medal zone with not ensemble but single model.\n\nSo then could you share your model and your generalization strategy?",
    "821132": "You could call the 43rd place solution a single model. They did ensemble multiple checkpoints though. They used EfficientNetB4 with added attention layers and wrote a full paper on it.\n\nDiscussion Post:\nhttps://www.kaggle.com/c/deepfake-detection-challenge/discussion/145841\n\nPaper:\nhttps://arxiv.org/pdf/2004.07676.pdf\n\nNotebook:\nhttps://www.kaggle.com/nicobonne/43-rank-ispl-ensamble-10-model\n\nGithub repo:\nhttps://github.com/polimi-ispl/icpr2020dfdc\n\nHope this helps!",
    "821816": "Thank you for you reply!\n\nEfficientnet is very remarkable for this competition besides widely known other-task.\nIt's very interesting they use just EfficientnetB4 for back-bone though train variants( attention / siamese) for ensembel.\n\nI guess they could have achieved better perfomance if they used more model(b5/b6...) for their ensemble.",
    "823713": "Literally in the last week of the competition, I tried a generalization technique which significantly improved my score. I was working on it through the week nights, and as such I ran out of time to generate proper ensemble solutions :(\nBut, since it was better than the others, I ended up making one of my two submissions a single model which scored 0.30848 in public LB and 0.48569 in private LB (it would place 47th place in private LB).\nThe only other submission I managed to train and submit within the deadline, was an ensemble of the first submission model, with another training run of the same model with a different random seed. The ensemble public LB was 0.30313 (66th place) and the private LB was 0.47007 ( 21st place ).\nI have to say that given the quick improvement in score for the simple two model ensemble, I'm quite annoyed that I didn't get a chance to prepare a proper ensemble.",
    "823884": "Congrats to your decent outcome though you didn't have a chances to test your ensemble!\n\nEnsemble was so effective also for this competition especially enhancing generalization power \nconsidering all experience shraed on disscussion board.\n\nSo, what kind of model did you use for your backbone? I'm curious about that.",
    "823908": "Thanks. I used EfficientNet B7 with input size 224",
    "824108": "Our team decided to go for just a single model. Maybe if we had an ensemble, we would have reached a higher place due to generalization. Single model is a effnet-b4\nSome tweaks: \n- Range of (heavy) augmentations used\n- A less risky submission with probabilities clipped between 0.15 and 0.9\n- Clever face detection and multiple actor tracking system",
    "824274": "wow...!",
    "824275": "Stunning result... You get generalized model without ensemble! Would you mind if you share the technique?!",
    "824280": "Default Input size for b7 is 600. \nDo you have any special reason to use input size 224?\n\nit's interesting \nbecause I thought larger input-size have benefit to detect blending boundary.",
    "824283": "Effnet worked superbly for this competition.\n\nIt's very nice to achiev high-rank with single model.\n\nWould you share multiple face tracking system in detail?",
    "825224": "Thanks! True that, effnets deemed to be more superior for us more than any other model. To the point I was speculating most teams were using them too.\n\nRegarding the multiple face tracking system - credits goes to my team-mate @yannael, so forgive me if there is some faults in my somewhat complicated description.\n\nBriefly, we track a person based on the bounding boxes of the faces. Our goal is to give every person a unique ID based on these boxes. We define a threshold distance of around ~100 pixels between the box coordinates. If that threshold is surpassed, a new ID box is created assuming there is a second person in the video. Then for each consecutive frame, we associate the new bounding box to their IDs based on the minimal distance to the last boxes from the previous frame.",
    "825484": "The original face sizes vary. I was concerned about up-scaling smallish faces too much and having them look fake due to upscaling artifacts, etc. I was trying to pick a size to minimize the average amount of up-scaling required. I tried some smaller sizes and some larger sizes and this size worked well.",
    "829790": "On the other hand I concerned about resizing to smaller size compared to original might destroy manipulation artifacts. I tested input size 256-640 and conclude to use 512 resolution.\nNow I read your idea and there might be pros and cons of either resizing larger or smaller."
  },
  "source": "meta"
}