{
  "id": 140358,
  "title": "What will be the efficient generalization strategy?",
  "url": "/competitions/deepfake-detection-challenge/discussion/140358",
  "author_name": "",
  "post_date": "2020-04-01T13:04:39.036387500Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Final score might be dependent on generalization power.</p>\n\n<p>One index to evaluate generalization power would be gap between LB and CV score.</p>\n\n<p>Half of the models in my ensemble was so overfitted with CV dataset. (CV 0.12 and LB 0.36)\nAnd rest of the model have a gap (CV 0.18 and LB 0.37 )</p>\n\n<p>Though I use small model ( less than,  &lt; 3million parameter, &lt; 10MB with 32bit float ) to avoid to overfit, it was easily overfitted to train dataset.\nSo I worry about private evaluation :)</p>\n\n<p>@harangdev wrote that increasing magnitude of augumentation in accordance with model size and face crop-size would alleviate overfitting on his kind article.</p>\n\n<p>BTW, in my case augumentation only decrease gap (between CV and LB) by 0.02.</p>\n\n<p>Do you think what would be a efficient generalization strategy?</p>\n\n<p>Use external data, Model capacity, Augumentation.... etc</p>\n\n<p>And what is the index for evaluate generalization other than CV / LB Scores?</p>",
  "messages": [
    {
      "id": "793980",
      "postDate": "04/01/2020 13:04:39",
      "content": "<p>Final score might be dependent on generalization power.</p>\n\n<p>One index to evaluate generalization power would be gap between LB and CV score.</p>\n\n<p>Half of the models in my ensemble was so overfitted with CV dataset. (CV 0.12 and LB 0.36)\nAnd rest of the model have a gap (CV 0.18 and LB 0.37 )</p>\n\n<p>Though I use small model ( less than,  &lt; 3million parameter, &lt; 10MB with 32bit float ) to avoid to overfit, it was easily overfitted to train dataset.\nSo I worry about private evaluation :)</p>\n\n<p>@harangdev wrote that increasing magnitude of augumentation in accordance with model size and face crop-size would alleviate overfitting on his kind article.</p>\n\n<p>BTW, in my case augumentation only decrease gap (between CV and LB) by 0.02.</p>\n\n<p>Do you think what would be a efficient generalization strategy?</p>\n\n<p>Use external data, Model capacity, Augumentation.... etc</p>\n\n<p>And what is the index for evaluate generalization other than CV / LB Scores?</p>",
      "rawMarkdown": "Final score might be dependent on generalization power.\n\nOne index to evaluate generalization power would be gap between LB and CV score.\n\nHalf of the models in my ensemble was so overfitted with CV dataset. (CV 0.12 and LB 0.36)\nAnd rest of the model have a gap (CV 0.18 and LB 0.37 )\n\nThough I use small model ( less than,  &lt; 3million parameter, &lt; 10MB with 32bit float ) to avoid to overfit, it was easily overfitted to train dataset.\nSo I worry about private evaluation :)\n\n@harangdev wrote that increasing magnitude of augumentation in accordance with model size and face crop-size would alleviate overfitting on his kind article.\n\nBTW, in my case augumentation only decrease gap (between CV and LB) by 0.02.\n\nDo you think what would be a efficient generalization strategy?\n\nUse external data, Model capacity, Augumentation.... etc\n\nAnd what is the index for evaluate generalization other than CV / LB Scores?",
      "votes": null
    },
    {
      "id": "794008",
      "postDate": "04/01/2020 13:30:05",
      "content": "<p>The answer to that will fully depend on private dataset and in what ways it differs from validation and train set.</p>",
      "rawMarkdown": "The answer to that will fully depend on private dataset and in what ways it differs from validation and train set.",
      "votes": null
    },
    {
      "id": "794023",
      "postDate": "04/01/2020 13:46:04",
      "content": "<p>Yes, and I hope boarder prepare fairly selected private datasets not biased.</p>",
      "rawMarkdown": "Yes, and I hope boarder prepare fairly selected private datasets not biased.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 794008,
      "author_name": "miguelpm",
      "author_url": "",
      "post_date": "04/01/2020 13:30:05",
      "content": "<p>The answer to that will fully depend on private dataset and in what ways it differs from validation and train set.</p>",
      "votes": null,
      "replies": [
        {
          "id": 794023,
          "author_name": "gwsong",
          "author_url": "",
          "post_date": "04/01/2020 13:46:04",
          "content": "<p>Yes, and I hope boarder prepare fairly selected private datasets not biased.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "793980": "Final score might be dependent on generalization power.\n\nOne index to evaluate generalization power would be gap between LB and CV score.\n\nHalf of the models in my ensemble was so overfitted with CV dataset. (CV 0.12 and LB 0.36)\nAnd rest of the model have a gap (CV 0.18 and LB 0.37 )\n\nThough I use small model ( less than,  &lt; 3million parameter, &lt; 10MB with 32bit float ) to avoid to overfit, it was easily overfitted to train dataset.\nSo I worry about private evaluation :)\n\n@harangdev wrote that increasing magnitude of augumentation in accordance with model size and face crop-size would alleviate overfitting on his kind article.\n\nBTW, in my case augumentation only decrease gap (between CV and LB) by 0.02.\n\nDo you think what would be a efficient generalization strategy?\n\nUse external data, Model capacity, Augumentation.... etc\n\nAnd what is the index for evaluate generalization other than CV / LB Scores?",
    "794008": "The answer to that will fully depend on private dataset and in what ways it differs from validation and train set.",
    "794023": "Yes, and I hope boarder prepare fairly selected private datasets not biased."
  },
  "source": "meta"
}