{
  "id": 142946,
  "title": "Is Separable Conv effective in DeepFake Detection? And why?",
  "url": "/competitions/deepfake-detection-challenge/discussion/142946",
  "author_name": "",
  "post_date": "2020-04-13T02:25:37.849903400Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>It seems that backbones with Separable Conv, like EfficientNet and Xception, outperform in this task. Does it implies that Separable Conv plays an important role? Maybe it's worth to make studies about it. </p>",
  "messages": [
    {
      "id": "805730",
      "postDate": "04/13/2020 02:25:37",
      "content": "<p>It seems that backbones with Separable Conv, like EfficientNet and Xception, outperform in this task. Does it implies that Separable Conv plays an important role? Maybe it's worth to make studies about it. </p>",
      "rawMarkdown": "It seems that backbones with Separable Conv, like EfficientNet and Xception, outperform in this task. Does it implies that Separable Conv plays an important role? Maybe it's worth to make studies about it.",
      "votes": null
    },
    {
      "id": "806052",
      "postDate": "04/13/2020 12:01:03",
      "content": "<p>One intuition is that it's simply a more efficient parameterization. Regular convs input every channel for every output, which gives you N^2 parameters (*kernel_size). \nUsing groups conserves the input coverage (every channels is used somewhere) but drastically reduces the number of parameters. \nThis allows to have a larger model (in the sense of deeper or wider) within your gpu and/or makes the learning much faster (larger batch size and less multiplications/operations)</p>\n\n<p>I wonder if anyone has tried it on the fully connected layers as well... ?</p>",
      "rawMarkdown": "One intuition is that it's simply a more efficient parameterization. Regular convs input every channel for every output, which gives you N^2 parameters (*kernel_size). \nUsing groups conserves the input coverage (every channels is used somewhere) but drastically reduces the number of parameters. \nThis allows to have a larger model (in the sense of deeper or wider) within your gpu and/or makes the learning much faster (larger batch size and less multiplications/operations)\n\nI wonder if anyone has tried it on the fully connected layers as well... ?",
      "votes": null
    },
    {
      "id": "806293",
      "postDate": "04/13/2020 15:23:07",
      "content": "<p>I did not have time to investigate this but i think the densenet will be the most effective at detecting deep fakes. My intuition comes from the IEEE camera model challenge, where in my experiments a single  densenet model outperformed ensembles of many other models. Why? I think in problems like camera model identification and deep fakes, the first few convolutional layers in the network are super important. The image features are so subtle, that the quality of the features learnt by the first few layers closest to the image input determine the outcome.  My theory is that densent does a better job at preserving these low level features.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F369212%2F0da0c38d01ebd696d008f1f44da692cd%2Fdensenet1.png?generation=1586791170428074&amp;alt=media\" alt=\"\"></p>\n\n<p>More evidence of this is the 2nd place write up where they said they used a u-net. Initially you might think the multitask learning for classification and segmentation is what gave them a boost. I think what gave them a boost is the similarity between the U-net and the densenet with the shortcut connections. THE U-NET PRESERVES low level features with .the short-cut connections.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F369212%2F240be1004df7e9b03ca515d223888bd6%2Fu-net-architecture.png?generation=1586792526289584&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I did not have time to investigate this but i think the densenet will be the most effective at detecting deep fakes. My intuition comes from the IEEE camera model challenge, where in my experiments a single  densenet model outperformed ensembles of many other models. Why? I think in problems like camera model identification and deep fakes, the first few convolutional layers in the network are super important. The image features are so subtle, that the quality of the features learnt by the first few layers closest to the image input determine the outcome.  My theory is that densent does a better job at preserving these low level features.![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F369212%2F0da0c38d01ebd696d008f1f44da692cd%2Fdensenet1.png?generation=1586791170428074&amp;alt=media)\n\nMore evidence of this is the 2nd place write up where they said they used a u-net. Initially you might think the multitask learning for classification and segmentation is what gave them a boost. I think what gave them a boost is the similarity between the U-net and the densenet with the shortcut connections. THE U-NET PRESERVES low level features with .the short-cut connections.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F369212%2F240be1004df7e9b03ca515d223888bd6%2Fu-net-architecture.png?generation=1586792526289584&amp;alt=media)",
      "votes": null
    },
    {
      "id": "806527",
      "postDate": "04/13/2020 19:45:04",
      "content": "<p>Isn't it effective in all kinds of tasks? With less parameters and same capacity, you can reduce the chance of overfitting. (I guess)</p>",
      "rawMarkdown": "Isn't it effective in all kinds of tasks? With less parameters and same capacity, you can reduce the chance of overfitting. (I guess)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 806052,
      "author_name": "hmendonca",
      "author_url": "",
      "post_date": "04/13/2020 12:01:03",
      "content": "<p>One intuition is that it's simply a more efficient parameterization. Regular convs input every channel for every output, which gives you N^2 parameters (*kernel_size). \nUsing groups conserves the input coverage (every channels is used somewhere) but drastically reduces the number of parameters. \nThis allows to have a larger model (in the sense of deeper or wider) within your gpu and/or makes the learning much faster (larger batch size and less multiplications/operations)</p>\n\n<p>I wonder if anyone has tried it on the fully connected layers as well... ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 806293,
      "author_name": "godaibo",
      "author_url": "",
      "post_date": "04/13/2020 15:23:07",
      "content": "<p>I did not have time to investigate this but i think the densenet will be the most effective at detecting deep fakes. My intuition comes from the IEEE camera model challenge, where in my experiments a single  densenet model outperformed ensembles of many other models. Why? I think in problems like camera model identification and deep fakes, the first few convolutional layers in the network are super important. The image features are so subtle, that the quality of the features learnt by the first few layers closest to the image input determine the outcome.  My theory is that densent does a better job at preserving these low level features.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F369212%2F0da0c38d01ebd696d008f1f44da692cd%2Fdensenet1.png?generation=1586791170428074&amp;alt=media\" alt=\"\"></p>\n\n<p>More evidence of this is the 2nd place write up where they said they used a u-net. Initially you might think the multitask learning for classification and segmentation is what gave them a boost. I think what gave them a boost is the similarity between the U-net and the densenet with the shortcut connections. THE U-NET PRESERVES low level features with .the short-cut connections.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F369212%2F240be1004df7e9b03ca515d223888bd6%2Fu-net-architecture.png?generation=1586792526289584&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 806527,
      "author_name": "yuanzhezhou",
      "author_url": "",
      "post_date": "04/13/2020 19:45:04",
      "content": "<p>Isn't it effective in all kinds of tasks? With less parameters and same capacity, you can reduce the chance of overfitting. (I guess)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "805730": "It seems that backbones with Separable Conv, like EfficientNet and Xception, outperform in this task. Does it implies that Separable Conv plays an important role? Maybe it's worth to make studies about it.",
    "806052": "One intuition is that it's simply a more efficient parameterization. Regular convs input every channel for every output, which gives you N^2 parameters (*kernel_size). \nUsing groups conserves the input coverage (every channels is used somewhere) but drastically reduces the number of parameters. \nThis allows to have a larger model (in the sense of deeper or wider) within your gpu and/or makes the learning much faster (larger batch size and less multiplications/operations)\n\nI wonder if anyone has tried it on the fully connected layers as well... ?",
    "806293": "I did not have time to investigate this but i think the densenet will be the most effective at detecting deep fakes. My intuition comes from the IEEE camera model challenge, where in my experiments a single  densenet model outperformed ensembles of many other models. Why? I think in problems like camera model identification and deep fakes, the first few convolutional layers in the network are super important. The image features are so subtle, that the quality of the features learnt by the first few layers closest to the image input determine the outcome.  My theory is that densent does a better job at preserving these low level features.![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F369212%2F0da0c38d01ebd696d008f1f44da692cd%2Fdensenet1.png?generation=1586791170428074&amp;alt=media)\n\nMore evidence of this is the 2nd place write up where they said they used a u-net. Initially you might think the multitask learning for classification and segmentation is what gave them a boost. I think what gave them a boost is the similarity between the U-net and the densenet with the shortcut connections. THE U-NET PRESERVES low level features with .the short-cut connections.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F369212%2F240be1004df7e9b03ca515d223888bd6%2Fu-net-architecture.png?generation=1586792526289584&amp;alt=media)",
    "806527": "Isn't it effective in all kinds of tasks? With less parameters and same capacity, you can reduce the chance of overfitting. (I guess)"
  },
  "source": "meta"
}