{
  "id": 134845,
  "title": "The magic of 0.69314",
  "url": "/competitions/deepfake-detection-challenge/discussion/134845",
  "author_name": "",
  "post_date": "2020-03-10T18:06:48.738491600Z",
  "votes": 1,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hello, kagglers.\nThe log-loss when all predictions are 0.5 on balanced dataset is ~0.69314, so this is just a baseline number.\nMy neural network stucks at this number(train log-loss) during training for a half of an epoch(~140000 images per one epoch). \nSince time is quite important resource here - do you think it is worth waiting and continue training or better to switch to another model?</p>",
  "messages": [
    {
      "id": "768381",
      "postDate": "03/10/2020 18:06:48",
      "content": "<p>Hello, kagglers.\nThe log-loss when all predictions are 0.5 on balanced dataset is ~0.69314, so this is just a baseline number.\nMy neural network stucks at this number(train log-loss) during training for a half of an epoch(~140000 images per one epoch). \nSince time is quite important resource here - do you think it is worth waiting and continue training or better to switch to another model?</p>",
      "rawMarkdown": "Hello, kagglers.\nThe log-loss when all predictions are 0.5 on balanced dataset is ~0.69314, so this is just a baseline number.\nMy neural network stucks at this number(train log-loss) during training for a half of an epoch(~140000 images per one epoch). \nSince time is quite important resource here - do you think it is worth waiting and continue training or better to switch to another model?",
      "votes": null
    },
    {
      "id": "768393",
      "postDate": "03/10/2020 18:31:43",
      "content": "<p>It is probably a more simple /complicated problem and not the model. Check your images. If you feed model with, for example, the same image with 0 and 1 labels, you will get the behavior you describe. </p>",
      "rawMarkdown": "It is probably a more simple /complicated problem and not the model. Check your images. If you feed model with, for example, the same image with 0 and 1 labels, you will get the behavior you describe.",
      "votes": null
    },
    {
      "id": "768529",
      "postDate": "03/10/2020 23:37:50",
      "content": "<p><a href=\"/vpaslay\">@vpaslay</a> It should be able to get under 0.69314 in ~200 samples(notice, not batches), and stabilize in ~500 samples. Otherwise, there's something wrong. Check your generator, images, model, and hyperparameters(such as optimizer, number of neurons for dense).\n<a href=\"/moshel\">@moshel</a> Sometimes, the model might actually be the problem. Once, Dense(64) is bad for us, but Dense(128) worked. </p>\n\n<p>Not sure in your case or other people's cases but its what we observed.</p>",
      "rawMarkdown": "vpaslay It should be able to get under 0.69314 in ~200 samples(notice, not batches), and stabilize in ~500 samples. Otherwise, there's something wrong. Check your generator, images, model, and hyperparameters(such as optimizer, number of neurons for dense).\n@moshel Sometimes, the model might actually be the problem. Once, Dense(64) is bad for us, but Dense(128) worked. \n\nNot sure in your case or other people's cases but its what we observed.",
      "votes": null
    },
    {
      "id": "768545",
      "postDate": "03/11/2020 00:00:07",
      "content": "<p>A common mistake is not to cast the image to float before dividing by 255. This results in mostly zero image... Hypothetically, of course. Not that it ever happened to me :) </p>",
      "rawMarkdown": "A common mistake is not to cast the image to float before dividing by 255. This results in mostly zero image... Hypothetically, of course. Not that it ever happened to me :)",
      "votes": null
    },
    {
      "id": "768581",
      "postDate": "03/11/2020 02:02:24",
      "content": "<p>I didn't know that\nI did a simple experiment:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3319985%2F142e73b7dbf8d78d133b095b6acd6e79%2FScreen%20Shot%202020-03-10%20at%207.00.51%20PM.png?generation=1583892068601215&amp;alt=media\" alt=\"\"></p>\n\n<p>Looks like we don't need that converstion on my system or kaggle's. IDK whether I understand you correctly. Please correct me if I'm wrong.</p>",
      "rawMarkdown": "I didn't know that\nI did a simple experiment:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3319985%2F142e73b7dbf8d78d133b095b6acd6e79%2FScreen%20Shot%202020-03-10%20at%207.00.51%20PM.png?generation=1583892068601215&amp;alt=media)\n\nLooks like we don't need that converstion on my system or kaggle's. IDK whether I understand you correctly. Please correct me if I'm wrong.",
      "votes": null
    },
    {
      "id": "768811",
      "postDate": "03/11/2020 08:00:06",
      "content": "<p>The input was completely correct. I figured that out when I have connected different model from keras applications - I have got pretty low train and validation loss(&lt;0.1). However, when I set efficientnet from here(<a href=\"https://github.com/qubvel/efficientnet\">https://github.com/qubvel/efficientnet</a>) as backbone - I get only this 0.69314. </p>\n\n<p><code>\nenet = efn.EfficientNetB7(weights='imagenet', pooling = 'avg') \nmodel = keras.Sequential([\n    enet,\n    keras.layers.Dense(1, activation='sigmoid')\n])\n</code></p>",
      "rawMarkdown": "The input was completely correct. I figured that out when I have connected different model from keras applications - I have got pretty low train and validation loss(&lt;0.1). However, when I set efficientnet from here(https://github.com/qubvel/efficientnet) as backbone - I get only this 0.69314. \n\n```\nenet = efn.EfficientNetB7(weights='imagenet', pooling = 'avg') \nmodel = keras.Sequential([\n    enet,\n    keras.layers.Dense(1, activation='sigmoid')\n])\n```",
      "votes": null
    },
    {
      "id": "768970",
      "postDate": "03/11/2020 11:52:30",
      "content": "<p><a href=\"/vpaslay\">@vpaslay</a> Umm, how can you that wrong lol (dont mind), an EffB7 is a total overkill, unless you give it face images of resolution close to 800x800 it does not even have a chance, just change the classifier to lower one b0-b4 is recommmened (depends so much on just the image input size).</p>",
      "rawMarkdown": "vpaslay Umm, how can you that wrong lol (dont mind), an EffB7 is a total overkill, unless you give it face images of resolution close to 800x800 it does not even have a chance, just change the classifier to lower one b0-b4 is recommmened (depends so much on just the image input size).",
      "votes": null
    },
    {
      "id": "768994",
      "postDate": "03/11/2020 12:16:04",
      "content": "<p>1) I am not using the DFDC dataset(but similar), so I have high-res images.\n2) The size of an image after the processing is 300x300.\n3) I have used NASNet with much greater success, though it has bigger number of parameters.\n4) EfficientNet has been tried by its authors on CIFAR-100 where the image size is 32x32.\nWhy do you think it is an overkill?</p>",
      "rawMarkdown": "1) I am not using the DFDC dataset(but similar), so I have high-res images.\n2) The size of an image after the processing is 300x300.\n3) I have used NASNet with much greater success, though it has bigger number of parameters.\n4) EfficientNet has been tried by its authors on CIFAR-100 where the image size is 32x32.\nWhy do you think it is an overkill?",
      "votes": null
    },
    {
      "id": "769168",
      "postDate": "03/11/2020 15:45:07",
      "content": "<p><a href=\"/vpaslay\">@vpaslay</a> \n1. According to the code, the default size for effb7 is 600x600. Its not saying you can't use effb7 for 300x300. But 300x300 only provide enough information for effb3, training it with effb7 is a waste of parameters, only makes it easier to overfit.\n2. I'm not sure what \"success with NASNet\" mean, if it only have a good CV, but bad LB(not in comparison to your previous submission, but with all others), it probably overfit with faces.\n3. Only effb0 have been used in CIFAR-100. </p>\n\n<p>Anyways, I recommend you to try effb3(300x300 is the default input size for effb3). Just our thoughts, if you don't agree, that's completely OK. Please correct me if I'm wrong. </p>",
      "rawMarkdown": "vpaslay \n1. According to the code, the default size for effb7 is 600x600. Its not saying you can't use effb7 for 300x300. But 300x300 only provide enough information for effb3, training it with effb7 is a waste of parameters, only makes it easier to overfit.\n2. I'm not sure what \"success with NASNet\" mean, if it only have a good CV, but bad LB(not in comparison to your previous submission, but with all others), it probably overfit with faces.\n3. Only effb0 have been used in CIFAR-100. \n\nAnyways, I recommend you to try effb3(300x300 is the default input size for effb3). Just our thoughts, if you don't agree, that's completely OK. Please correct me if I'm wrong.",
      "votes": null
    },
    {
      "id": "769198",
      "postDate": "03/11/2020 16:22:52",
      "content": "<p><a href=\"/unkownhihi\">@unkownhihi</a>\nI will surely try, it is a good suggestion. What package would you recommend?\nAccording to the <a href=\"https://arxiv.org/pdf/1905.11946.pdf\">paper</a>(table 5) the authors trained effB7-imagenet on CIFAR-10 as well and got better results than effb0.\nI have submitted only 0.5 predictions until now, and I am discussing the behavior on the dataset uploaded by Bojan Tunguz. By the way NASNet works great on hold-out part of data(not validation data).</p>",
      "rawMarkdown": "unkownhihi\nI will surely try, it is a good suggestion. What package would you recommend?\nAccording to the [paper](https://arxiv.org/pdf/1905.11946.pdf)(table 5) the authors trained effB7-imagenet on CIFAR-10 as well and got better results than effb0.\nI have submitted only 0.5 predictions until now, and I am discussing the behavior on the dataset uploaded by Bojan Tunguz. By the way NASNet works great on hold-out part of data(not validation data).",
      "votes": null
    },
    {
      "id": "769219",
      "postDate": "03/11/2020 16:43:45",
      "content": "<p>for pytorch, you can use <a href=\"https://pypi.org/project/efficientnet-pytorch/\">link</a>\nfor keras you can use <a href=\"https://pypi.org/project/efficientnet/\">link</a>\nI understand why NASNet works well on Bojan Tunguz's dataset. It has 1 million images, while this dataset has ~100k.</p>",
      "rawMarkdown": "for pytorch, you can use [link](https://pypi.org/project/efficientnet-pytorch/)\nfor keras you can use [link](https://pypi.org/project/efficientnet/)\nI understand why NASNet works well on Bojan Tunguz's dataset. It has 1 million images, while this dataset has ~100k.",
      "votes": null
    },
    {
      "id": "770112",
      "postDate": "03/12/2020 15:17:27",
      "content": "<p>I was stuck during my training for a long time but things got much better after I've set weight decay around 0.01. My newbie mistake... :) \nAt least I've learned the importance of avoiding strong overfit in training data.\nIt seems EfficientNet can ovefit very quickly. You can see training loss going down by a lot while validation loss barely moves it.</p>",
      "rawMarkdown": "I was stuck during my training for a long time but things got much better after I've set weight decay around 0.01. My newbie mistake... :) \nAt least I've learned the importance of avoiding strong overfit in training data.\nIt seems EfficientNet can ovefit very quickly. You can see training loss going down by a lot while validation loss barely moves it.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 768393,
      "author_name": "moshel",
      "author_url": "",
      "post_date": "03/10/2020 18:31:43",
      "content": "<p>It is probably a more simple /complicated problem and not the model. Check your images. If you feed model with, for example, the same image with 0 and 1 labels, you will get the behavior you describe. </p>",
      "votes": null,
      "replies": [
        {
          "id": 768811,
          "author_name": "vpaslay",
          "author_url": "",
          "post_date": "03/11/2020 08:00:06",
          "content": "<p>The input was completely correct. I figured that out when I have connected different model from keras applications - I have got pretty low train and validation loss(&lt;0.1). However, when I set efficientnet from here(<a href=\"https://github.com/qubvel/efficientnet\">https://github.com/qubvel/efficientnet</a>) as backbone - I get only this 0.69314. </p>\n\n<p><code>\nenet = efn.EfficientNetB7(weights='imagenet', pooling = 'avg') \nmodel = keras.Sequential([\n    enet,\n    keras.layers.Dense(1, activation='sigmoid')\n])\n</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 768970,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "03/11/2020 11:52:30",
          "content": "<p><a href=\"/vpaslay\">@vpaslay</a> Umm, how can you that wrong lol (dont mind), an EffB7 is a total overkill, unless you give it face images of resolution close to 800x800 it does not even have a chance, just change the classifier to lower one b0-b4 is recommmened (depends so much on just the image input size).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 768994,
          "author_name": "vpaslay",
          "author_url": "",
          "post_date": "03/11/2020 12:16:04",
          "content": "<p>1) I am not using the DFDC dataset(but similar), so I have high-res images.\n2) The size of an image after the processing is 300x300.\n3) I have used NASNet with much greater success, though it has bigger number of parameters.\n4) EfficientNet has been tried by its authors on CIFAR-100 where the image size is 32x32.\nWhy do you think it is an overkill?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 769168,
          "author_name": "unkownhihi",
          "author_url": "",
          "post_date": "03/11/2020 15:45:07",
          "content": "<p><a href=\"/vpaslay\">@vpaslay</a> \n1. According to the code, the default size for effb7 is 600x600. Its not saying you can't use effb7 for 300x300. But 300x300 only provide enough information for effb3, training it with effb7 is a waste of parameters, only makes it easier to overfit.\n2. I'm not sure what \"success with NASNet\" mean, if it only have a good CV, but bad LB(not in comparison to your previous submission, but with all others), it probably overfit with faces.\n3. Only effb0 have been used in CIFAR-100. </p>\n\n<p>Anyways, I recommend you to try effb3(300x300 is the default input size for effb3). Just our thoughts, if you don't agree, that's completely OK. Please correct me if I'm wrong. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 769198,
          "author_name": "vpaslay",
          "author_url": "",
          "post_date": "03/11/2020 16:22:52",
          "content": "<p><a href=\"/unkownhihi\">@unkownhihi</a>\nI will surely try, it is a good suggestion. What package would you recommend?\nAccording to the <a href=\"https://arxiv.org/pdf/1905.11946.pdf\">paper</a>(table 5) the authors trained effB7-imagenet on CIFAR-10 as well and got better results than effb0.\nI have submitted only 0.5 predictions until now, and I am discussing the behavior on the dataset uploaded by Bojan Tunguz. By the way NASNet works great on hold-out part of data(not validation data).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 769219,
          "author_name": "unkownhihi",
          "author_url": "",
          "post_date": "03/11/2020 16:43:45",
          "content": "<p>for pytorch, you can use <a href=\"https://pypi.org/project/efficientnet-pytorch/\">link</a>\nfor keras you can use <a href=\"https://pypi.org/project/efficientnet/\">link</a>\nI understand why NASNet works well on Bojan Tunguz's dataset. It has 1 million images, while this dataset has ~100k.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 768529,
      "author_name": "unkownhihi",
      "author_url": "",
      "post_date": "03/10/2020 23:37:50",
      "content": "<p><a href=\"/vpaslay\">@vpaslay</a> It should be able to get under 0.69314 in ~200 samples(notice, not batches), and stabilize in ~500 samples. Otherwise, there's something wrong. Check your generator, images, model, and hyperparameters(such as optimizer, number of neurons for dense).\n<a href=\"/moshel\">@moshel</a> Sometimes, the model might actually be the problem. Once, Dense(64) is bad for us, but Dense(128) worked. </p>\n\n<p>Not sure in your case or other people's cases but its what we observed.</p>",
      "votes": null,
      "replies": [
        {
          "id": 768545,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "03/11/2020 00:00:07",
          "content": "<p>A common mistake is not to cast the image to float before dividing by 255. This results in mostly zero image... Hypothetically, of course. Not that it ever happened to me :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 768581,
          "author_name": "unkownhihi",
          "author_url": "",
          "post_date": "03/11/2020 02:02:24",
          "content": "<p>I didn't know that\nI did a simple experiment:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3319985%2F142e73b7dbf8d78d133b095b6acd6e79%2FScreen%20Shot%202020-03-10%20at%207.00.51%20PM.png?generation=1583892068601215&amp;alt=media\" alt=\"\"></p>\n\n<p>Looks like we don't need that converstion on my system or kaggle's. IDK whether I understand you correctly. Please correct me if I'm wrong.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 770112,
      "author_name": "mont3z",
      "author_url": "",
      "post_date": "03/12/2020 15:17:27",
      "content": "<p>I was stuck during my training for a long time but things got much better after I've set weight decay around 0.01. My newbie mistake... :) \nAt least I've learned the importance of avoiding strong overfit in training data.\nIt seems EfficientNet can ovefit very quickly. You can see training loss going down by a lot while validation loss barely moves it.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "768381": "Hello, kagglers.\nThe log-loss when all predictions are 0.5 on balanced dataset is ~0.69314, so this is just a baseline number.\nMy neural network stucks at this number(train log-loss) during training for a half of an epoch(~140000 images per one epoch). \nSince time is quite important resource here - do you think it is worth waiting and continue training or better to switch to another model?",
    "768393": "It is probably a more simple /complicated problem and not the model. Check your images. If you feed model with, for example, the same image with 0 and 1 labels, you will get the behavior you describe.",
    "768529": "vpaslay It should be able to get under 0.69314 in ~200 samples(notice, not batches), and stabilize in ~500 samples. Otherwise, there's something wrong. Check your generator, images, model, and hyperparameters(such as optimizer, number of neurons for dense).\n@moshel Sometimes, the model might actually be the problem. Once, Dense(64) is bad for us, but Dense(128) worked. \n\nNot sure in your case or other people's cases but its what we observed.",
    "768545": "A common mistake is not to cast the image to float before dividing by 255. This results in mostly zero image... Hypothetically, of course. Not that it ever happened to me :)",
    "768581": "I didn't know that\nI did a simple experiment:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3319985%2F142e73b7dbf8d78d133b095b6acd6e79%2FScreen%20Shot%202020-03-10%20at%207.00.51%20PM.png?generation=1583892068601215&amp;alt=media)\n\nLooks like we don't need that converstion on my system or kaggle's. IDK whether I understand you correctly. Please correct me if I'm wrong.",
    "768811": "The input was completely correct. I figured that out when I have connected different model from keras applications - I have got pretty low train and validation loss(&lt;0.1). However, when I set efficientnet from here(https://github.com/qubvel/efficientnet) as backbone - I get only this 0.69314. \n\n```\nenet = efn.EfficientNetB7(weights='imagenet', pooling = 'avg') \nmodel = keras.Sequential([\n    enet,\n    keras.layers.Dense(1, activation='sigmoid')\n])\n```",
    "768970": "vpaslay Umm, how can you that wrong lol (dont mind), an EffB7 is a total overkill, unless you give it face images of resolution close to 800x800 it does not even have a chance, just change the classifier to lower one b0-b4 is recommmened (depends so much on just the image input size).",
    "768994": "1) I am not using the DFDC dataset(but similar), so I have high-res images.\n2) The size of an image after the processing is 300x300.\n3) I have used NASNet with much greater success, though it has bigger number of parameters.\n4) EfficientNet has been tried by its authors on CIFAR-100 where the image size is 32x32.\nWhy do you think it is an overkill?",
    "769168": "vpaslay \n1. According to the code, the default size for effb7 is 600x600. Its not saying you can't use effb7 for 300x300. But 300x300 only provide enough information for effb3, training it with effb7 is a waste of parameters, only makes it easier to overfit.\n2. I'm not sure what \"success with NASNet\" mean, if it only have a good CV, but bad LB(not in comparison to your previous submission, but with all others), it probably overfit with faces.\n3. Only effb0 have been used in CIFAR-100. \n\nAnyways, I recommend you to try effb3(300x300 is the default input size for effb3). Just our thoughts, if you don't agree, that's completely OK. Please correct me if I'm wrong.",
    "769198": "unkownhihi\nI will surely try, it is a good suggestion. What package would you recommend?\nAccording to the [paper](https://arxiv.org/pdf/1905.11946.pdf)(table 5) the authors trained effB7-imagenet on CIFAR-10 as well and got better results than effb0.\nI have submitted only 0.5 predictions until now, and I am discussing the behavior on the dataset uploaded by Bojan Tunguz. By the way NASNet works great on hold-out part of data(not validation data).",
    "769219": "for pytorch, you can use [link](https://pypi.org/project/efficientnet-pytorch/)\nfor keras you can use [link](https://pypi.org/project/efficientnet/)\nI understand why NASNet works well on Bojan Tunguz's dataset. It has 1 million images, while this dataset has ~100k.",
    "770112": "I was stuck during my training for a long time but things got much better after I've set weight decay around 0.01. My newbie mistake... :) \nAt least I've learned the importance of avoiding strong overfit in training data.\nIt seems EfficientNet can ovefit very quickly. You can see training loss going down by a lot while validation loss barely moves it."
  },
  "source": "meta"
}