{
  "id": 131209,
  "title": "transfer learning vs model from scratch",
  "url": "/competitions/deepfake-detection-challenge/discussion/131209",
  "author_name": "",
  "post_date": "2020-02-18T21:21:45.123116400Z",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Playing a bit with a few models. So far my experience has been not so successful in the use of pre-trained models.. I have tried VGG16 from Keras, initially with only the top dense layers to learn and a successive fine tune of the last 2 conv layers.... training loss score never breached the 0.69 no skill score.</p>\n\n<p>Same training data, I have made up a much smaller and simple conv network and trained it from scratch, and training loss has easily reached 0.5.. i haven't tried it yet to submit this one.</p>\n\n<p>So I am just wondering if the top scoring models are mostly making use of pre-trained models or building brand new ones?</p>",
  "messages": [
    {
      "id": "749688",
      "postDate": "02/18/2020 21:21:45",
      "content": "<p>Playing a bit with a few models. So far my experience has been not so successful in the use of pre-trained models.. I have tried VGG16 from Keras, initially with only the top dense layers to learn and a successive fine tune of the last 2 conv layers.... training loss score never breached the 0.69 no skill score.</p>\n\n<p>Same training data, I have made up a much smaller and simple conv network and trained it from scratch, and training loss has easily reached 0.5.. i haven't tried it yet to submit this one.</p>\n\n<p>So I am just wondering if the top scoring models are mostly making use of pre-trained models or building brand new ones?</p>",
      "rawMarkdown": "Playing a bit with a few models. So far my experience has been not so successful in the use of pre-trained models.. I have tried VGG16 from Keras, initially with only the top dense layers to learn and a successive fine tune of the last 2 conv layers.... training loss score never breached the 0.69 no skill score.\n\nSame training data, I have made up a much smaller and simple conv network and trained it from scratch, and training loss has easily reached 0.5.. i haven't tried it yet to submit this one.\n\nSo I am just wondering if the top scoring models are mostly making use of pre-trained models or building brand new ones?",
      "votes": null
    },
    {
      "id": "749696",
      "postDate": "02/18/2020 21:27:23",
      "content": "<p><a href=\"/antonio1979\">@antonio1979</a> you can have a look at the top public kernels/notebooks and see what they are doing ;)</p>",
      "rawMarkdown": "antonio1979 you can have a look at the top public kernels/notebooks and see what they are doing ;)",
      "votes": null
    },
    {
      "id": "749706",
      "postDate": "02/18/2020 21:38:44",
      "content": "<p>Also, try to generate a good dataset, it makes a lot of difference. Hints are everywhere in the discussions!</p>",
      "rawMarkdown": "Also, try to generate a good dataset, it makes a lot of difference. Hints are everywhere in the discussions!",
      "votes": null
    },
    {
      "id": "750049",
      "postDate": "02/19/2020 05:01:03",
      "content": "<p>VGG16 is way too small!</p>",
      "rawMarkdown": "VGG16 is way too small!",
      "votes": null
    },
    {
      "id": "750287",
      "postDate": "02/19/2020 09:04:34",
      "content": "<p>is it? excluding the top dense layers, it's a good 15m parameters and it's not able to learn behind the 0.69 score. What i find interesting and can't explain is that my custom 3m parameters model is able to score better than VGG16</p>",
      "rawMarkdown": "is it? excluding the top dense layers, it's a good 15m parameters and it's not able to learn behind the 0.69 score. What i find interesting and can't explain is that my custom 3m parameters model is able to score better than VGG16",
      "votes": null
    },
    {
      "id": "750422",
      "postDate": "02/19/2020 11:02:09",
      "content": "<blockquote>\n  <p>my custom 3m parameters model is able to score better than VGG16</p>\n</blockquote>\n\n<p>You only mentioned that your training loss went down. All this means is that your smaller model is learning something, not whether it learns the right thing. So saying \"it is able to score better\" may not be correct.</p>\n\n<p>That said, there is no reason why VGG16 shouldn't learn. Perhaps there is a mistake in your training code, e.g. you did not set the layers to be learnable or your LR is too small, etc.</p>",
      "rawMarkdown": "&gt; my custom 3m parameters model is able to score better than VGG16\n\nYou only mentioned that your training loss went down. All this means is that your smaller model is learning something, not whether it learns the right thing. So saying \"it is able to score better\" may not be correct.\n\nThat said, there is no reason why VGG16 shouldn't learn. Perhaps there is a mistake in your training code, e.g. you did not set the layers to be learnable or your LR is too small, etc.",
      "votes": null
    },
    {
      "id": "750457",
      "postDate": "02/19/2020 11:52:46",
      "content": "<p>thanks for the feedback... i see... i'll give another try with VGG16 then, or do you think it's a too small a model too?</p>",
      "rawMarkdown": "thanks for the feedback... i see... i'll give another try with VGG16 then, or do you think it's a too small a model too?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 749696,
      "author_name": "hmendonca",
      "author_url": "",
      "post_date": "02/18/2020 21:27:23",
      "content": "<p><a href=\"/antonio1979\">@antonio1979</a> you can have a look at the top public kernels/notebooks and see what they are doing ;)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 749706,
      "author_name": "debanga",
      "author_url": "",
      "post_date": "02/18/2020 21:38:44",
      "content": "<p>Also, try to generate a good dataset, it makes a lot of difference. Hints are everywhere in the discussions!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 750049,
      "author_name": "unkownhihi",
      "author_url": "",
      "post_date": "02/19/2020 05:01:03",
      "content": "<p>VGG16 is way too small!</p>",
      "votes": null,
      "replies": [
        {
          "id": 750287,
          "author_name": "antonio1979",
          "author_url": "",
          "post_date": "02/19/2020 09:04:34",
          "content": "<p>is it? excluding the top dense layers, it's a good 15m parameters and it's not able to learn behind the 0.69 score. What i find interesting and can't explain is that my custom 3m parameters model is able to score better than VGG16</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 750422,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "02/19/2020 11:02:09",
          "content": "<blockquote>\n  <p>my custom 3m parameters model is able to score better than VGG16</p>\n</blockquote>\n\n<p>You only mentioned that your training loss went down. All this means is that your smaller model is learning something, not whether it learns the right thing. So saying \"it is able to score better\" may not be correct.</p>\n\n<p>That said, there is no reason why VGG16 shouldn't learn. Perhaps there is a mistake in your training code, e.g. you did not set the layers to be learnable or your LR is too small, etc.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 750457,
          "author_name": "antonio1979",
          "author_url": "",
          "post_date": "02/19/2020 11:52:46",
          "content": "<p>thanks for the feedback... i see... i'll give another try with VGG16 then, or do you think it's a too small a model too?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "749688": "Playing a bit with a few models. So far my experience has been not so successful in the use of pre-trained models.. I have tried VGG16 from Keras, initially with only the top dense layers to learn and a successive fine tune of the last 2 conv layers.... training loss score never breached the 0.69 no skill score.\n\nSame training data, I have made up a much smaller and simple conv network and trained it from scratch, and training loss has easily reached 0.5.. i haven't tried it yet to submit this one.\n\nSo I am just wondering if the top scoring models are mostly making use of pre-trained models or building brand new ones?",
    "749696": "antonio1979 you can have a look at the top public kernels/notebooks and see what they are doing ;)",
    "749706": "Also, try to generate a good dataset, it makes a lot of difference. Hints are everywhere in the discussions!",
    "750049": "VGG16 is way too small!",
    "750287": "is it? excluding the top dense layers, it's a good 15m parameters and it's not able to learn behind the 0.69 score. What i find interesting and can't explain is that my custom 3m parameters model is able to score better than VGG16",
    "750422": "&gt; my custom 3m parameters model is able to score better than VGG16\n\nYou only mentioned that your training loss went down. All this means is that your smaller model is learning something, not whether it learns the right thing. So saying \"it is able to score better\" may not be correct.\n\nThat said, there is no reason why VGG16 shouldn't learn. Perhaps there is a mistake in your training code, e.g. you did not set the layers to be learnable or your LR is too small, etc.",
    "750457": "thanks for the feedback... i see... i'll give another try with VGG16 then, or do you think it's a too small a model too?"
  },
  "source": "meta"
}