{
  "id": 132885,
  "title": "Score difference between test and LB results",
  "url": "/competitions/deepfake-detection-challenge/discussion/132885",
  "author_name": "Daytona DING",
  "post_date": "2020-02-28T12:44:14.745000",
  "votes": 2,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I split the data sets as following: 0-40 training set, 41-46 validation set, 47-49 test set. The model gave me ~0.38 on the test set, but only achieved 0.69 on LB. Do you have any idea what may cause this provided the code is correct?</p>",
  "messages": [
    {
      "id": 759142,
      "postDate": "2020-02-28T16:29:18.093Z",
      "content": "<p>Have you tried to undersample your dataset to solve the problem of unbalanced data? I found this kernel very helpful: <a href=\"https://www.kaggle.com/rafjaa/resampling-strategies-for-imbalanced-datasets#t5\">https://www.kaggle.com/rafjaa/resampling-strategies-for-imbalanced-datasets#t5</a></p>",
      "rawMarkdown": "Have you tried to undersample your dataset to solve the problem of unbalanced data? I found this kernel very helpful: https://www.kaggle.com/rafjaa/resampling-strategies-for-imbalanced-datasets#t5",
      "votes": 1
    },
    {
      "id": 759112,
      "postDate": "2020-02-28T15:45:18.773Z",
      "content": "<p>What about a memory leak that takes place only on the 4000 videos (say after #500...) ???\nDon't forget the \"submit\" code mostly don't \"crash\", but keeps running no matter what - even if a model.load('') has failed or something similar.</p>",
      "rawMarkdown": "What about a memory leak that takes place only on the 4000 videos (say after #500...) ???\nDon't forget the \"submit\" code mostly don't \"crash\", but keeps running no matter what - even if a model.load('') has failed or something similar.",
      "votes": 1
    },
    {
      "id": 759109,
      "postDate": "2020-02-28T15:42:23.483Z",
      "content": "<p>Funny fact I got exactly the same numbers. 0.38 becomes .69 (val set is 8, 13, 18, 23 etc...). Data pollution is not an issue and has been taken care of.</p>\n\n<p>Takes 8 hours to compute... as if everything was fine... there's definitely some assumption you and I probably make about the data that is completely incorrect. I'm re-reading my code like I'm a total retard. Still not getting it. Same code runs fine on my side. Even checked tensorflow and numpy versions issues...</p>\n\n<p>Actually, it's the second (fully different) model that I design with similar results. Tried to  \"try/execpt\" some critical portion with almost no success. Both models tried had very reasonable overfit on local tests.</p>\n\n<p>Frame numbers ? -&gt; would have lessen the compute time\nVideo resolution ? -&gt; would have lessen (or increased) the compute time\nWeird number of faces per video ? Why not...\nDifferent Codec used -&gt; inconsistency in RGB outputs ?</p>\n\n<p>Any idea ?</p>",
      "rawMarkdown": "Funny fact I got exactly the same numbers. 0.38 becomes .69 (val set is 8, 13, 18, 23 etc...). Data pollution is not an issue and has been taken care of.\n\nTakes 8 hours to compute... as if everything was fine... there's definitely some assumption you and I probably make about the data that is completely incorrect. I'm re-reading my code like I'm a total retard. Still not getting it. Same code runs fine on my side. Even checked tensorflow and numpy versions issues...\n\nActually, it's the second (fully different) model that I design with similar results. Tried to  \"try/execpt\" some critical portion with almost no success. Both models tried had very reasonable overfit on local tests.\n\nFrame numbers ? -&gt; would have lessen the compute time\nVideo resolution ? -&gt; would have lessen (or increased) the compute time\nWeird number of faces per video ? Why not...\nDifferent Codec used -&gt; inconsistency in RGB outputs ?\n\nAny idea ?",
      "votes": 1
    },
    {
      "id": 759430,
      "postDate": "2020-02-29T02:51:30.987Z",
      "content": "<p>You need to calibrate the \"probabilities\" (they are not true probabilities) produced by the network. Neural networks are notorious at being too confident in the prediction particularly if there is bias in the training data. If you eliminate all bias in the training set it is less of a problem but in this challenge it is hard to do that without hand picking every frame. Cross entropy loss will significantly penalise confident predictions that are wrong so it is better to be less confident than more confident when faced with training bias. There are different approaches to calibrating NNs. Here is a simple one to look into.</p>\n\n<p><a href=\"https://geoffpleiss.com/nn_calibration\">https://geoffpleiss.com/nn_calibration</a></p>",
      "rawMarkdown": "You need to calibrate the \"probabilities\" (they are not true probabilities) produced by the network. Neural networks are notorious at being too confident in the prediction particularly if there is bias in the training data. If you eliminate all bias in the training set it is less of a problem but in this challenge it is hard to do that without hand picking every frame. Cross entropy loss will significantly penalise confident predictions that are wrong so it is better to be less confident than more confident when faced with training bias. There are different approaches to calibrating NNs. Here is a simple one to look into.\n\nhttps://geoffpleiss.com/nn_calibration",
      "votes": 2,
      "replies": [
        {
          "id": 759804,
          "postDate": "2020-02-29T13:37:06.260Z",
          "content": "<p>Thanks for sharing this idea ! This is pretty much what I'm implementing right now. Using prior knowledge to calibrate the typical output curve. I usually end up with .8 to 1.2 power-curve and a very small added bias. I'm still trying to figure out why I need to do this to my model ? Usually these models are meant to be log-loss effective... Thanks to these DFDC, I figured out they need some (a lot of) tweaking.\nFunny fact :  train you model for another 0.1 epoch, and you have to complete change the power-curve and bias... a kind of random noise within the last layers ? Reducing learning speed ? But the current speed is Ok for the upper part of the model ?</p>",
          "rawMarkdown": "Thanks for sharing this idea ! This is pretty much what I'm implementing right now. Using prior knowledge to calibrate the typical output curve. I usually end up with .8 to 1.2 power-curve and a very small added bias. I'm still trying to figure out why I need to do this to my model ? Usually these models are meant to be log-loss effective... Thanks to these DFDC, I figured out they need some (a lot of) tweaking.\nFunny fact :  train you model for another 0.1 epoch, and you have to complete change the power-curve and bias... a kind of random noise within the last layers ? Reducing learning speed ? But the current speed is Ok for the upper part of the model ?"
        },
        {
          "id": 759836,
          "postDate": "2020-02-29T14:34:45.133Z",
          "content": "<p>IMO the best approach is to avoid bias in data so that different types of observations are evenly balanced given one model. Failing that you can use ensembles to help address the uncertainty introduced by bias because N+1 predictions averaged will eliminate more uncertainty than one prediction. The theory is similar to random forests. Also, rather than a power curve you might want to look at Isotonic Regression if you stick to one model. My preference would be ensembles. I think you will find with large batch sizes the calibration is more volatile. I suggest smaller batch sizes (16 to 32). Btw my current ranking does not reflect these ideas yet :)</p>",
          "rawMarkdown": "IMO the best approach is to avoid bias in data so that different types of observations are evenly balanced given one model. Failing that you can use ensembles to help address the uncertainty introduced by bias because N+1 predictions averaged will eliminate more uncertainty than one prediction. The theory is similar to random forests. Also, rather than a power curve you might want to look at Isotonic Regression if you stick to one model. My preference would be ensembles. I think you will find with large batch sizes the calibration is more volatile. I suggest smaller batch sizes (16 to 32). Btw my current ranking does not reflect these ideas yet :)"
        },
        {
          "id": 760080,
          "postDate": "2020-02-29T20:47:19.623Z",
          "content": "<p>Learnt something new today. Thanks mate!</p>",
          "rawMarkdown": "Learnt something new today. Thanks mate!"
        }
      ]
    },
    {
      "id": 758999,
      "postDate": "2020-02-28T12:44:14.747Z",
      "content": "<p>I split the data sets as following: 0-40 training set, 41-46 validation set, 47-49 test set. The model gave me ~0.38 on the test set, but only achieved 0.69 on LB. Do you have any idea what may cause this provided the code is correct?</p>",
      "rawMarkdown": " I split the data sets as following: 0-40 training set, 41-46 validation set, 47-49 test set. The model gave me ~0.38 on the test set, but only achieved 0.69 on LB. Do you have any idea what may cause this provided the code is correct?",
      "votes": 2
    },
    {
      "id": 759612,
      "postDate": "2020-02-29T09:00:40.683Z",
      "content": "<p>wow</p>",
      "rawMarkdown": "wow"
    },
    {
      "id": 759018,
      "postDate": "2020-02-28T13:03:51.697Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 759142,
      "author_name": "Mont3z Claro5",
      "author_url": "",
      "post_date": "2020-02-28T16:29:18.093000",
      "content": "<p>Have you tried to undersample your dataset to solve the problem of unbalanced data? I found this kernel very helpful: <a href=\"https://www.kaggle.com/rafjaa/resampling-strategies-for-imbalanced-datasets#t5\">https://www.kaggle.com/rafjaa/resampling-strategies-for-imbalanced-datasets#t5</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 759112,
      "author_name": "Simon Caby",
      "author_url": "",
      "post_date": "2020-02-28T15:45:18.773000",
      "content": "<p>What about a memory leak that takes place only on the 4000 videos (say after #500...) ???\nDon't forget the \"submit\" code mostly don't \"crash\", but keeps running no matter what - even if a model.load('') has failed or something similar.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 759109,
      "author_name": "Simon Caby",
      "author_url": "",
      "post_date": "2020-02-28T15:42:23.483000",
      "content": "<p>Funny fact I got exactly the same numbers. 0.38 becomes .69 (val set is 8, 13, 18, 23 etc...). Data pollution is not an issue and has been taken care of.</p>\n\n<p>Takes 8 hours to compute... as if everything was fine... there's definitely some assumption you and I probably make about the data that is completely incorrect. I'm re-reading my code like I'm a total retard. Still not getting it. Same code runs fine on my side. Even checked tensorflow and numpy versions issues...</p>\n\n<p>Actually, it's the second (fully different) model that I design with similar results. Tried to  \"try/execpt\" some critical portion with almost no success. Both models tried had very reasonable overfit on local tests.</p>\n\n<p>Frame numbers ? -&gt; would have lessen the compute time\nVideo resolution ? -&gt; would have lessen (or increased) the compute time\nWeird number of faces per video ? Why not...\nDifferent Codec used -&gt; inconsistency in RGB outputs ?</p>\n\n<p>Any idea ?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 759430,
      "author_name": "maralski",
      "author_url": "",
      "post_date": "2020-02-29T02:51:30.987000",
      "content": "<p>You need to calibrate the \"probabilities\" (they are not true probabilities) produced by the network. Neural networks are notorious at being too confident in the prediction particularly if there is bias in the training data. If you eliminate all bias in the training set it is less of a problem but in this challenge it is hard to do that without hand picking every frame. Cross entropy loss will significantly penalise confident predictions that are wrong so it is better to be less confident than more confident when faced with training bias. There are different approaches to calibrating NNs. Here is a simple one to look into.</p>\n\n<p><a href=\"https://geoffpleiss.com/nn_calibration\">https://geoffpleiss.com/nn_calibration</a></p>",
      "votes": 2,
      "replies": [
        {
          "id": 759804,
          "author_name": "Simon Caby",
          "author_url": "",
          "post_date": "2020-02-29T13:37:06.260000",
          "content": "<p>Thanks for sharing this idea ! This is pretty much what I'm implementing right now. Using prior knowledge to calibrate the typical output curve. I usually end up with .8 to 1.2 power-curve and a very small added bias. I'm still trying to figure out why I need to do this to my model ? Usually these models are meant to be log-loss effective... Thanks to these DFDC, I figured out they need some (a lot of) tweaking.\nFunny fact :  train you model for another 0.1 epoch, and you have to complete change the power-curve and bias... a kind of random noise within the last layers ? Reducing learning speed ? But the current speed is Ok for the upper part of the model ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 759836,
          "author_name": "maralski",
          "author_url": "",
          "post_date": "2020-02-29T14:34:45.133000",
          "content": "<p>IMO the best approach is to avoid bias in data so that different types of observations are evenly balanced given one model. Failing that you can use ensembles to help address the uncertainty introduced by bias because N+1 predictions averaged will eliminate more uncertainty than one prediction. The theory is similar to random forests. Also, rather than a power curve you might want to look at Isotonic Regression if you stick to one model. My preference would be ensembles. I think you will find with large batch sizes the calibration is more volatile. I suggest smaller batch sizes (16 to 32). Btw my current ranking does not reflect these ideas yet :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 760080,
          "author_name": "Akash",
          "author_url": "",
          "post_date": "2020-02-29T20:47:19.623000",
          "content": "<p>Learnt something new today. Thanks mate!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 759612,
      "author_name": "Eva Mondal",
      "author_url": "",
      "post_date": "2020-02-29T09:00:40.683000",
      "content": "<p>wow</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 759018,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-28T13:03:51.697000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "759142": "Have you tried to undersample your dataset to solve the problem of unbalanced data? I found this kernel very helpful: https://www.kaggle.com/rafjaa/resampling-strategies-for-imbalanced-datasets#t5",
    "759112": "What about a memory leak that takes place only on the 4000 videos (say after #500...) ???\nDon't forget the \"submit\" code mostly don't \"crash\", but keeps running no matter what - even if a model.load('') has failed or something similar.",
    "759109": "Funny fact I got exactly the same numbers. 0.38 becomes .69 (val set is 8, 13, 18, 23 etc...). Data pollution is not an issue and has been taken care of.\n\nTakes 8 hours to compute... as if everything was fine... there's definitely some assumption you and I probably make about the data that is completely incorrect. I'm re-reading my code like I'm a total retard. Still not getting it. Same code runs fine on my side. Even checked tensorflow and numpy versions issues...\n\nActually, it's the second (fully different) model that I design with similar results. Tried to  \"try/execpt\" some critical portion with almost no success. Both models tried had very reasonable overfit on local tests.\n\nFrame numbers ? -&gt; would have lessen the compute time\nVideo resolution ? -&gt; would have lessen (or increased) the compute time\nWeird number of faces per video ? Why not...\nDifferent Codec used -&gt; inconsistency in RGB outputs ?\n\nAny idea ?",
    "759430": "You need to calibrate the \"probabilities\" (they are not true probabilities) produced by the network. Neural networks are notorious at being too confident in the prediction particularly if there is bias in the training data. If you eliminate all bias in the training set it is less of a problem but in this challenge it is hard to do that without hand picking every frame. Cross entropy loss will significantly penalise confident predictions that are wrong so it is better to be less confident than more confident when faced with training bias. There are different approaches to calibrating NNs. Here is a simple one to look into.\n\nhttps://geoffpleiss.com/nn_calibration",
    "758999": " I split the data sets as following: 0-40 training set, 41-46 validation set, 47-49 test set. The model gave me ~0.38 on the test set, but only achieved 0.69 on LB. Do you have any idea what may cause this provided the code is correct?",
    "759612": "wow",
    "759018": ""
  }
}