{
  "id": 130915,
  "title": "CV by actors?",
  "url": "/competitions/deepfake-detection-challenge/discussion/130915",
  "author_name": "",
  "post_date": "2020-02-17T05:55:11.312318300Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Based on some discussions, I was wondering if any one has tried to:</p>\n\n<ol>\n<li>Label each video with an actor id (using clustering or similar techniques as done <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/129832\">here</a>)</li>\n<li>Create a CV schema where each fold contains approximately the same percentage of actors. </li>\n</ol>\n\n<p>Apparently, some folders have more actors diversity (i.e. much more different actors) than some where you will find only few. </p>\n\n<p>Any thoughts on this?</p>",
  "messages": [
    {
      "id": "748027",
      "postDate": "02/17/2020 05:55:11",
      "content": "<p>Based on some discussions, I was wondering if any one has tried to:</p>\n\n<ol>\n<li>Label each video with an actor id (using clustering or similar techniques as done <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/129832\">here</a>)</li>\n<li>Create a CV schema where each fold contains approximately the same percentage of actors. </li>\n</ol>\n\n<p>Apparently, some folders have more actors diversity (i.e. much more different actors) than some where you will find only few. </p>\n\n<p>Any thoughts on this?</p>",
      "rawMarkdown": "Based on some discussions, I was wondering if any one has tried to:\n\n1. Label each video with an actor id (using clustering or similar techniques as done [here](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/129832))\n2. Create a CV schema where each fold contains approximately the same percentage of actors. \n\nApparently, some folders have more actors diversity (i.e. much more different actors) than some where you will find only few. \n\nAny thoughts on this?",
      "votes": null
    },
    {
      "id": "748697",
      "postDate": "02/17/2020 21:19:22",
      "content": "<p>So, the first point is not valid because we are not able to cluster faces in a very good way and second point is dependent on the first one, using folders 40-50 track the leaderboard well, in my case the leaderboard score will be 0.10 - 0.15 higher than my cv.</p>",
      "rawMarkdown": "So, the first point is not valid because we are not able to cluster faces in a very good way and second point is dependent on the first one, using folders 40-50 track the leaderboard well, in my case the leaderboard score will be 0.10 - 0.15 higher than my cv.",
      "votes": null
    },
    {
      "id": "748719",
      "postDate": "02/17/2020 22:09:33",
      "content": "<p>just an unrelated comment, k fold is really unneeded here... you have heaps of data as it is. focus on finding a relatively consistent cv like <a href=\"/harshitsheoran\">@harshitsheoran</a> said. I am still struggling with this, but getting there.</p>",
      "rawMarkdown": "just an unrelated comment, k fold is really unneeded here... you have heaps of data as it is. focus on finding a relatively consistent cv like @harshitsheoran said. I am still struggling with this, but getting there.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 748697,
      "author_name": "harshitsheoran",
      "author_url": "",
      "post_date": "02/17/2020 21:19:22",
      "content": "<p>So, the first point is not valid because we are not able to cluster faces in a very good way and second point is dependent on the first one, using folders 40-50 track the leaderboard well, in my case the leaderboard score will be 0.10 - 0.15 higher than my cv.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 748719,
      "author_name": "moshel",
      "author_url": "",
      "post_date": "02/17/2020 22:09:33",
      "content": "<p>just an unrelated comment, k fold is really unneeded here... you have heaps of data as it is. focus on finding a relatively consistent cv like <a href=\"/harshitsheoran\">@harshitsheoran</a> said. I am still struggling with this, but getting there.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "748027": "Based on some discussions, I was wondering if any one has tried to:\n\n1. Label each video with an actor id (using clustering or similar techniques as done [here](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/129832))\n2. Create a CV schema where each fold contains approximately the same percentage of actors. \n\nApparently, some folders have more actors diversity (i.e. much more different actors) than some where you will find only few. \n\nAny thoughts on this?",
    "748697": "So, the first point is not valid because we are not able to cluster faces in a very good way and second point is dependent on the first one, using folders 40-50 track the leaderboard well, in my case the leaderboard score will be 0.10 - 0.15 higher than my cv.",
    "748719": "just an unrelated comment, k fold is really unneeded here... you have heaps of data as it is. focus on finding a relatively consistent cv like @harshitsheoran said. I am still struggling with this, but getting there."
  },
  "source": "meta"
}