{
  "id": 136985,
  "title": "Validation loss is not enough to choose a model ",
  "url": "/competitions/deepfake-detection-challenge/discussion/136985",
  "author_name": "",
  "post_date": "2020-03-18T17:57:14.909711900Z",
  "votes": 3,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I've submitted three different models with similar cross validation loss (I'm validating on 11k frames in different folders than the one training on) between 0.33-0.30, but each model on LB got different score, only one performed better (0.4x) vs 0.55+ </p>\n\n<p>I'm pretty much new to deep learning.. is there a good metric to indicate if a model may generalize well on unseen data beside cross validation loss ? I've been trying to use ROC_AUC but these three models had ~0.93 score and still only one was 'good'.</p>\n\n<p>Cheers</p>",
  "messages": [
    {
      "id": "778766",
      "postDate": "03/18/2020 17:57:14",
      "content": "<p>I've submitted three different models with similar cross validation loss (I'm validating on 11k frames in different folders than the one training on) between 0.33-0.30, but each model on LB got different score, only one performed better (0.4x) vs 0.55+ </p>\n\n<p>I'm pretty much new to deep learning.. is there a good metric to indicate if a model may generalize well on unseen data beside cross validation loss ? I've been trying to use ROC_AUC but these three models had ~0.93 score and still only one was 'good'.</p>\n\n<p>Cheers</p>",
      "rawMarkdown": "I've submitted three different models with similar cross validation loss (I'm validating on 11k frames in different folders than the one training on) between 0.33-0.30, but each model on LB got different score, only one performed better (0.4x) vs 0.55+ \n\nI'm pretty much new to deep learning.. is there a good metric to indicate if a model may generalize well on unseen data beside cross validation loss ? I've been trying to use ROC_AUC but these three models had ~0.93 score and still only one was 'good'.\n\nCheers",
      "votes": null
    },
    {
      "id": "778864",
      "postDate": "03/18/2020 20:10:31",
      "content": "<p>We have this problem too. Our score used to correspond pretty well(about 2x), but once the validation score got down to 0.14, it became bad. Our potential theory is that there's more wild/difficult faces that is used to calculate LB. So now the problem is generalization. We are currently trying to add more data.</p>",
      "rawMarkdown": "We have this problem too. Our score used to correspond pretty well(about 2x), but once the validation score got down to 0.14, it became bad. Our potential theory is that there's more wild/difficult faces that is used to calculate LB. So now the problem is generalization. We are currently trying to add more data.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 778864,
      "author_name": "unkownhihi",
      "author_url": "",
      "post_date": "03/18/2020 20:10:31",
      "content": "<p>We have this problem too. Our score used to correspond pretty well(about 2x), but once the validation score got down to 0.14, it became bad. Our potential theory is that there's more wild/difficult faces that is used to calculate LB. So now the problem is generalization. We are currently trying to add more data.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "778766": "I've submitted three different models with similar cross validation loss (I'm validating on 11k frames in different folders than the one training on) between 0.33-0.30, but each model on LB got different score, only one performed better (0.4x) vs 0.55+ \n\nI'm pretty much new to deep learning.. is there a good metric to indicate if a model may generalize well on unseen data beside cross validation loss ? I've been trying to use ROC_AUC but these three models had ~0.93 score and still only one was 'good'.\n\nCheers",
    "778864": "We have this problem too. Our score used to correspond pretty well(about 2x), but once the validation score got down to 0.14, it became bad. Our potential theory is that there's more wild/difficult faces that is used to calculate LB. So now the problem is generalization. We are currently trying to add more data."
  },
  "source": "meta"
}