{
  "id": 145713,
  "title": "AMA Interview with Yifan Xie | Competition Winning Team | Chai Time Data Science Show",
  "url": "/competitions/deepfake-detection-challenge/discussion/145713",
  "author_name": "",
  "post_date": "2020-04-24T08:14:05.144798Z",
  "votes": 10,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi Everyone!</p>\n\n<p>A big congratulations to the dream team on the 1st position finish on Private LB, Since I've been very fortunate to have interviewed <a href=\"/titericz\">@titericz</a> <a href=\"https://www.youtube.com/watch?v=MpYeDKw8EOg\">link to video</a>, <a href=\"https://anchor.fm/chaitimedatascience/episodes/Kaggle-Legend-Gilberto-Titericz--Giba--Former-1--Data-Science--Kaggle-ea9vdn/a-a1b5bid\">Link to audio</a> and <a href=\"/anokas\">@anokas</a> <a href=\"https://www.youtube.com/watch?v=maR9ibJ2r7g\">Link to video</a> and <a href=\"https://anchor.fm/chaitimedatascience/episodes/Anokas-Mikel-Bober-Irizar--Becoming-The-Youngest-Kaggle-Grandmaster--ML-For-Japanese-Literature--Kaggle-eanr0n/a-a1e8ad0\">Link to audio</a> earlier;</p>\n\n<p>So I reached out to <a href=\"/yifanxie\">@yifanxie</a>, who very kindly agreed to allow an AMA.</p>\n\n<p>I invite all/any questions as replies to this thread. </p>\n\n<p>About:</p>\n\n<p>This will be released on the <a href=\"http://chaitimedatascience.com\">Chai Time Data Science Podcast</a>, available both as <a href=\"https://www.youtube.com/playlist?list=PLLvvXm0q8zUbiNdoIazGzlENMXvZ9bd3x\">video</a>, <a href=\"https://anchor.fm/chaitimedatascience\">audio</a> and <a href=\"http://sanyambhutani.com/tag/chaitimedatascience\">Blog</a></p>\n\n<p>The interview date is yet to be scheduled, meanwhile, I'll keep collecting all questions. </p>\n\n<p>Thank You!\nAnd Thanks to Yifan!</p>\n\n<p>Edit: The interview is scheduled for 18/6/20, I'll keep collecting all questions till 12PM GMT 18/6. </p>\n\n<p>Unfortunately, There was a change in final LB, <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/157983\">read here</a>.</p>",
  "messages": [
    {
      "id": "818936",
      "postDate": "04/24/2020 08:14:05",
      "content": "<p>Hi Everyone!</p>\n\n<p>A big congratulations to the dream team on the 1st position finish on Private LB, Since I've been very fortunate to have interviewed <a href=\"/titericz\">@titericz</a> <a href=\"https://www.youtube.com/watch?v=MpYeDKw8EOg\">link to video</a>, <a href=\"https://anchor.fm/chaitimedatascience/episodes/Kaggle-Legend-Gilberto-Titericz--Giba--Former-1--Data-Science--Kaggle-ea9vdn/a-a1b5bid\">Link to audio</a> and <a href=\"/anokas\">@anokas</a> <a href=\"https://www.youtube.com/watch?v=maR9ibJ2r7g\">Link to video</a> and <a href=\"https://anchor.fm/chaitimedatascience/episodes/Anokas-Mikel-Bober-Irizar--Becoming-The-Youngest-Kaggle-Grandmaster--ML-For-Japanese-Literature--Kaggle-eanr0n/a-a1e8ad0\">Link to audio</a> earlier;</p>\n\n<p>So I reached out to <a href=\"/yifanxie\">@yifanxie</a>, who very kindly agreed to allow an AMA.</p>\n\n<p>I invite all/any questions as replies to this thread. </p>\n\n<p>About:</p>\n\n<p>This will be released on the <a href=\"http://chaitimedatascience.com\">Chai Time Data Science Podcast</a>, available both as <a href=\"https://www.youtube.com/playlist?list=PLLvvXm0q8zUbiNdoIazGzlENMXvZ9bd3x\">video</a>, <a href=\"https://anchor.fm/chaitimedatascience\">audio</a> and <a href=\"http://sanyambhutani.com/tag/chaitimedatascience\">Blog</a></p>\n\n<p>The interview date is yet to be scheduled, meanwhile, I'll keep collecting all questions. </p>\n\n<p>Thank You!\nAnd Thanks to Yifan!</p>\n\n<p>Edit: The interview is scheduled for 18/6/20, I'll keep collecting all questions till 12PM GMT 18/6. </p>\n\n<p>Unfortunately, There was a change in final LB, <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/157983\">read here</a>.</p>",
      "rawMarkdown": "Hi Everyone!\n\nA big congratulations to the dream team on the 1st position finish on Private LB, Since I've been very fortunate to have interviewed @titericz [link to video](https://www.youtube.com/watch?v=MpYeDKw8EOg), [Link to audio](https://anchor.fm/chaitimedatascience/episodes/Kaggle-Legend-Gilberto-Titericz--Giba--Former-1--Data-Science--Kaggle-ea9vdn/a-a1b5bid) and @anokas [Link to video](https://www.youtube.com/watch?v=maR9ibJ2r7g) and [Link to audio](https://anchor.fm/chaitimedatascience/episodes/Anokas-Mikel-Bober-Irizar--Becoming-The-Youngest-Kaggle-Grandmaster--ML-For-Japanese-Literature--Kaggle-eanr0n/a-a1e8ad0) earlier;\n\nSo I reached out to @yifanxie, who very kindly agreed to allow an AMA.\n\nI invite all/any questions as replies to this thread. \n\nAbout:\n\nThis will be released on the [Chai Time Data Science Podcast](http://chaitimedatascience.com), available both as [video](https://www.youtube.com/playlist?list=PLLvvXm0q8zUbiNdoIazGzlENMXvZ9bd3x), [audio](https://anchor.fm/chaitimedatascience) and [Blog](http://sanyambhutani.com/tag/chaitimedatascience)\n\nThe interview date is yet to be scheduled, meanwhile, I'll keep collecting all questions. \n\nThank You!\nAnd Thanks to Yifan!\n\nEdit: The interview is scheduled for 18/6/20, I'll keep collecting all questions till 12PM GMT 18/6. \n\nUnfortunately, There was a change in final LB, [read here](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/157983).",
      "votes": null
    },
    {
      "id": "818955",
      "postDate": "04/24/2020 08:32:31",
      "content": "<p>Congratulations! Did you feel confident that you would get a big boost in your score in the shake up? Did you feel that you'd overfit less than others, or thought you might generalise particularly well?</p>\n\n<p>Did you train on videos outside of the official training data?</p>",
      "rawMarkdown": "Congratulations! Did you feel confident that you would get a big boost in your score in the shake up? Did you feel that you'd overfit less than others, or thought you might generalise particularly well?\n\nDid you train on videos outside of the official training data?",
      "votes": null
    },
    {
      "id": "818972",
      "postDate": "04/24/2020 08:52:44",
      "content": "<p>Thanks a lot for your interview!</p>\n\n<ol>\n<li><p>It's the overfitting &amp; generalization issue what most of the teams have during competition undergoing.\nDoes AFAR-team have any special strategy for getting generalization power?</p></li>\n<li><p>If the team used ensemble methods, which model and how many models used?</p></li>\n<li><p>Did the team evaluate their model for another fake-methods?\nWhat's the hardest point of their model in determining whether the video is fake or not?\nIs it dependent on video characteristic / specific fake-methods ?</p></li>\n<li><p>Your model performance is only 0.42 logloss ( It's great for sure but not confident to detect )\n Do you have any idea to improve the detection performance?</p></li>\n</ol>\n\n<p>Thank you again :)</p>",
      "rawMarkdown": "Thanks a lot for your interview!\n\n1. It's the overfitting &amp; generalization issue what most of the teams have during competition undergoing.\n   Does AFAR-team have any special strategy for getting generalization power?\n\n2. If the team used ensemble methods, which model and how many models used?\n\n3. Did the team evaluate their model for another fake-methods?\n    What's the hardest point of their model in determining whether the video is fake or not?\n    Is it dependent on video characteristic / specific fake-methods ?\n\n4. Your model performance is only 0.42 logloss ( It's great for sure but not confident to detect )\n     Do you have any idea to improve the detection performance?\n\nThank you again :)",
      "votes": null
    },
    {
      "id": "818993",
      "postDate": "04/24/2020 09:12:06",
      "content": "<p>Will you spend all that prize money on new GPUs? 😄 </p>",
      "rawMarkdown": "Will you spend all that prize money on new GPUs? 😄",
      "votes": null
    },
    {
      "id": "819002",
      "postDate": "04/24/2020 09:16:59",
      "content": "<p>I'm actually kind of curious how the process of building this team worked, and how you all worked together. Having all masters and grandmasters obviously gives you edge over people with less experience. Do you believe that was indeed the case? </p>\n\n<p>Sanyam called you a \"dream team\" and I'm just wondering if working together with these highly skilled people really was a \"dream\" compared to other teams you may have been on? (Sometimes highly skilled people also have very strong opinions and that can make working together more difficult.)</p>",
      "rawMarkdown": "I'm actually kind of curious how the process of building this team worked, and how you all worked together. Having all masters and grandmasters obviously gives you edge over people with less experience. Do you believe that was indeed the case? \n\nSanyam called you a \"dream team\" and I'm just wondering if working together with these highly skilled people really was a \"dream\" compared to other teams you may have been on? (Sometimes highly skilled people also have very strong opinions and that can make working together more difficult.)",
      "votes": null
    },
    {
      "id": "819187",
      "postDate": "04/24/2020 12:14:01",
      "content": "<p>thanks all for your questions :) \nI am completely overrun with messages and the working day, but will get back to you in time, either via here or via Sanyam with his podcast session  </p>\n\n<p>will do proper solution sharing as well</p>",
      "rawMarkdown": "thanks all for your questions :) \nI am completely overrun with messages and the working day, but will get back to you in time, either via here or via Sanyam with his podcast session  \n\nwill do proper solution sharing as well",
      "votes": null
    },
    {
      "id": "821771",
      "postDate": "04/26/2020 12:11:41",
      "content": "<ol>\n<li>Please rank the methods/strategies/tricks you used that give you best improvement in term of Public LB score (please provide the logLoss before and after using those methods if possible).</li>\n<li>Please do let us know what methods doesn't work.</li>\n</ol>",
      "rawMarkdown": "1. Please rank the methods/strategies/tricks you used that give you best improvement in term of Public LB score (please provide the logLoss before and after using those methods if possible).\n2. Please do let us know what methods doesn't work.",
      "votes": null
    },
    {
      "id": "824932",
      "postDate": "04/28/2020 17:11:40",
      "content": "<p>Apart from syntax, tactical codes and programming caveats- what are your top underlying principles/ meta behind solving tough data science/ even other types of problems in general?</p>",
      "rawMarkdown": "Apart from syntax, tactical codes and programming caveats- what are your top underlying principles/ meta behind solving tough data science/ even other types of problems in general?",
      "votes": null
    },
    {
      "id": "827695",
      "postDate": "04/30/2020 13:17:39",
      "content": "<p>Congratulation once again! <a href=\"/yifanxie\">@yifanxie</a> </p>",
      "rawMarkdown": "Congratulation once again! @yifanxie",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 818955,
      "author_name": "jamesphoward",
      "author_url": "",
      "post_date": "04/24/2020 08:32:31",
      "content": "<p>Congratulations! Did you feel confident that you would get a big boost in your score in the shake up? Did you feel that you'd overfit less than others, or thought you might generalise particularly well?</p>\n\n<p>Did you train on videos outside of the official training data?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 818972,
      "author_name": "gwsong",
      "author_url": "",
      "post_date": "04/24/2020 08:52:44",
      "content": "<p>Thanks a lot for your interview!</p>\n\n<ol>\n<li><p>It's the overfitting &amp; generalization issue what most of the teams have during competition undergoing.\nDoes AFAR-team have any special strategy for getting generalization power?</p></li>\n<li><p>If the team used ensemble methods, which model and how many models used?</p></li>\n<li><p>Did the team evaluate their model for another fake-methods?\nWhat's the hardest point of their model in determining whether the video is fake or not?\nIs it dependent on video characteristic / specific fake-methods ?</p></li>\n<li><p>Your model performance is only 0.42 logloss ( It's great for sure but not confident to detect )\n Do you have any idea to improve the detection performance?</p></li>\n</ol>\n\n<p>Thank you again :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 818993,
      "author_name": "humananalog",
      "author_url": "",
      "post_date": "04/24/2020 09:12:06",
      "content": "<p>Will you spend all that prize money on new GPUs? 😄 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 819002,
      "author_name": "humananalog",
      "author_url": "",
      "post_date": "04/24/2020 09:16:59",
      "content": "<p>I'm actually kind of curious how the process of building this team worked, and how you all worked together. Having all masters and grandmasters obviously gives you edge over people with less experience. Do you believe that was indeed the case? </p>\n\n<p>Sanyam called you a \"dream team\" and I'm just wondering if working together with these highly skilled people really was a \"dream\" compared to other teams you may have been on? (Sometimes highly skilled people also have very strong opinions and that can make working together more difficult.)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 819187,
      "author_name": "yifanxie",
      "author_url": "",
      "post_date": "04/24/2020 12:14:01",
      "content": "<p>thanks all for your questions :) \nI am completely overrun with messages and the working day, but will get back to you in time, either via here or via Sanyam with his podcast session  </p>\n\n<p>will do proper solution sharing as well</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 821771,
      "author_name": "chewkokwahibrainai",
      "author_url": "",
      "post_date": "04/26/2020 12:11:41",
      "content": "<ol>\n<li>Please rank the methods/strategies/tricks you used that give you best improvement in term of Public LB score (please provide the logLoss before and after using those methods if possible).</li>\n<li>Please do let us know what methods doesn't work.</li>\n</ol>",
      "votes": null,
      "replies": []
    },
    {
      "id": 824932,
      "author_name": "fannnu",
      "author_url": "",
      "post_date": "04/28/2020 17:11:40",
      "content": "<p>Apart from syntax, tactical codes and programming caveats- what are your top underlying principles/ meta behind solving tough data science/ even other types of problems in general?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 827695,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "04/30/2020 13:17:39",
      "content": "<p>Congratulation once again! <a href=\"/yifanxie\">@yifanxie</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "818936": "Hi Everyone!\n\nA big congratulations to the dream team on the 1st position finish on Private LB, Since I've been very fortunate to have interviewed @titericz [link to video](https://www.youtube.com/watch?v=MpYeDKw8EOg), [Link to audio](https://anchor.fm/chaitimedatascience/episodes/Kaggle-Legend-Gilberto-Titericz--Giba--Former-1--Data-Science--Kaggle-ea9vdn/a-a1b5bid) and @anokas [Link to video](https://www.youtube.com/watch?v=maR9ibJ2r7g) and [Link to audio](https://anchor.fm/chaitimedatascience/episodes/Anokas-Mikel-Bober-Irizar--Becoming-The-Youngest-Kaggle-Grandmaster--ML-For-Japanese-Literature--Kaggle-eanr0n/a-a1e8ad0) earlier;\n\nSo I reached out to @yifanxie, who very kindly agreed to allow an AMA.\n\nI invite all/any questions as replies to this thread. \n\nAbout:\n\nThis will be released on the [Chai Time Data Science Podcast](http://chaitimedatascience.com), available both as [video](https://www.youtube.com/playlist?list=PLLvvXm0q8zUbiNdoIazGzlENMXvZ9bd3x), [audio](https://anchor.fm/chaitimedatascience) and [Blog](http://sanyambhutani.com/tag/chaitimedatascience)\n\nThe interview date is yet to be scheduled, meanwhile, I'll keep collecting all questions. \n\nThank You!\nAnd Thanks to Yifan!\n\nEdit: The interview is scheduled for 18/6/20, I'll keep collecting all questions till 12PM GMT 18/6. \n\nUnfortunately, There was a change in final LB, [read here](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/157983).",
    "818955": "Congratulations! Did you feel confident that you would get a big boost in your score in the shake up? Did you feel that you'd overfit less than others, or thought you might generalise particularly well?\n\nDid you train on videos outside of the official training data?",
    "818972": "Thanks a lot for your interview!\n\n1. It's the overfitting &amp; generalization issue what most of the teams have during competition undergoing.\n   Does AFAR-team have any special strategy for getting generalization power?\n\n2. If the team used ensemble methods, which model and how many models used?\n\n3. Did the team evaluate their model for another fake-methods?\n    What's the hardest point of their model in determining whether the video is fake or not?\n    Is it dependent on video characteristic / specific fake-methods ?\n\n4. Your model performance is only 0.42 logloss ( It's great for sure but not confident to detect )\n     Do you have any idea to improve the detection performance?\n\nThank you again :)",
    "818993": "Will you spend all that prize money on new GPUs? 😄",
    "819002": "I'm actually kind of curious how the process of building this team worked, and how you all worked together. Having all masters and grandmasters obviously gives you edge over people with less experience. Do you believe that was indeed the case? \n\nSanyam called you a \"dream team\" and I'm just wondering if working together with these highly skilled people really was a \"dream\" compared to other teams you may have been on? (Sometimes highly skilled people also have very strong opinions and that can make working together more difficult.)",
    "819187": "thanks all for your questions :) \nI am completely overrun with messages and the working day, but will get back to you in time, either via here or via Sanyam with his podcast session  \n\nwill do proper solution sharing as well",
    "821771": "1. Please rank the methods/strategies/tricks you used that give you best improvement in term of Public LB score (please provide the logLoss before and after using those methods if possible).\n2. Please do let us know what methods doesn't work.",
    "824932": "Apart from syntax, tactical codes and programming caveats- what are your top underlying principles/ meta behind solving tough data science/ even other types of problems in general?",
    "827695": "Congratulation once again! @yifanxie"
  },
  "source": "meta"
}