{
  "id": 127695,
  "title": "Big shakeup at top",
  "url": "/competitions/deepfake-detection-challenge/discussion/127695",
  "author_name": "",
  "post_date": "2020-01-25T23:16:56.005812300Z",
  "votes": 6,
  "comment_count": 38,
  "views": 0,
  "content": "<p>18 submissions with exactly 0.46788 in the last 10 hours, several first time submissions. How does that happen?</p>",
  "messages": [
    {
      "id": "729223",
      "postDate": "01/25/2020 23:16:56",
      "content": "<p>18 submissions with exactly 0.46788 in the last 10 hours, several first time submissions. How does that happen?</p>",
      "rawMarkdown": "18 submissions with exactly 0.46788 in the last 10 hours, several first time submissions. How does that happen?",
      "votes": null
    },
    {
      "id": "729224",
      "postDate": "01/25/2020 23:17:51",
      "content": "<p>Check <a href=\"https://www.kaggle.com/humananalog/inference-demo\">here</a></p>",
      "rawMarkdown": "Check [here](https://www.kaggle.com/humananalog/inference-demo)",
      "votes": null
    },
    {
      "id": "729225",
      "postDate": "01/25/2020 23:19:11",
      "content": "<p>Yes, just saw it thanks. I missed it because it wasn't flagged on the LB. Bummer to have to compete with a silver kernel.</p>",
      "rawMarkdown": "Yes, just saw it thanks. I missed it because it wasn't flagged on the LB. Bummer to have to compete with a silver kernel.",
      "votes": null
    },
    {
      "id": "729238",
      "postDate": "01/25/2020 23:50:09",
      "content": "<p>I'm the author of that kernel. 0.46788 really isn't that great of score on a logistic loss metric, so don't feel too discouraged...</p>\n\n<p>I suggest using this kernel as a baseline. If you use the same process with your own model (i.e. extract 17 or so frames from each video, do a prediction on each frame, take the mean) but your model scores less than 0.46x, it's a sign you need to improve your training data. </p>\n\n<p>As I said in the description, this uses a very simple ResNeXt50 model. It was only trained for about 7 epochs on ~30000 training images (split evenly between real/fake). I didn't spend any time on getting the best possible model design or training procedure. Each of my improvements on the leaderboard was done by improving my set of training images.</p>\n\n<p>So if you take the same approach as this kernel and your score is worse, spend some more time improving your dataset. Keep doing that and I guarantee you'll get a score below 0.46x. 😄 </p>",
      "rawMarkdown": "I'm the author of that kernel. 0.46788 really isn't that great of score on a logistic loss metric, so don't feel too discouraged...\n\nI suggest using this kernel as a baseline. If you use the same process with your own model (i.e. extract 17 or so frames from each video, do a prediction on each frame, take the mean) but your model scores less than 0.46x, it's a sign you need to improve your training data. \n\nAs I said in the description, this uses a very simple ResNeXt50 model. It was only trained for about 7 epochs on ~30000 training images (split evenly between real/fake). I didn't spend any time on getting the best possible model design or training procedure. Each of my improvements on the leaderboard was done by improving my set of training images.\n\nSo if you take the same approach as this kernel and your score is worse, spend some more time improving your dataset. Keep doing that and I guarantee you'll get a score below 0.46x. 😄",
      "votes": null
    },
    {
      "id": "729239",
      "postDate": "01/25/2020 23:51:33",
      "content": "<p>Considering there are 2 months left, I don't think 0.46788 would give any medals at the end. </p>",
      "rawMarkdown": "Considering there are 2 months left, I don't think 0.46788 would give any medals at the end.",
      "votes": null
    },
    {
      "id": "729257",
      "postDate": "01/26/2020 00:31:52",
      "content": "<p>Been in a lot of competitions over the past two years - I don't recall that I have ever seen a public shared kernel - at the two month point - be in the metal.</p>\n\n<p>I was very worried that the prize money would keep folks from posting good kernels.  Until Human Analog posted his, all the shared kernels had a leader board score that was pretty much the same that you can get with a random number generated submission.</p>\n\n<p>Hoping that his share will kick a few more loose rocks out of the mountain - pretty hard for the community to help in the creation of a worthy product when we are keeping our cards hidden.</p>",
      "rawMarkdown": "Been in a lot of competitions over the past two years - I don't recall that I have ever seen a public shared kernel - at the two month point - be in the metal.\n\nI was very worried that the prize money would keep folks from posting good kernels.  Until Human Analog posted his, all the shared kernels had a leader board score that was pretty much the same that you can get with a random number generated submission.\n\nHoping that his share will kick a few more loose rocks out of the mountain - pretty hard for the community to help in the creation of a worthy product when we are keeping our cards hidden.",
      "votes": null
    },
    {
      "id": "729347",
      "postDate": "01/26/2020 04:34:30",
      "content": "<p>Actually I don't think 0.46788 is a good baseline. It's too high in LB.\nPS. Tips and a good pipeline(like yours) are already enough for starters. But a trained model made the LB useless and it reminded me of the leak.</p>",
      "rawMarkdown": "Actually I don't think 0.46788 is a good baseline. It's too high in LB.\nPS. Tips and a good pipeline(like yours) are already enough for starters. But a trained model made the LB useless and it reminded me of the leak.",
      "votes": null
    },
    {
      "id": "729595",
      "postDate": "01/26/2020 12:31:09",
      "content": "<p>Agreed, people who copy this kernel immediately beat 90% of entrants who will have done an order of magnitude more work. It's against the spirit of Kaggle.</p>",
      "rawMarkdown": "Agreed, people who copy this kernel immediately beat 90% of entrants who will have done an order of magnitude more work. It's against the spirit of Kaggle.",
      "votes": null
    },
    {
      "id": "729684",
      "postDate": "01/26/2020 14:21:15",
      "content": "<p>To me, the spirit of Kaggle is about learning to get better at doing data science together.</p>\n\n<p>People who just run high-scoring kernels to get a position on the leaderboard, is a side-effect of allowing people to share kernels. It may not be ideal, but it's also not a big deal. A couple of weeks ago, a lot of people shot up on the leaderboards because they all ran the script that exploited the leakage. It happens. The only way to avoid that is to disable public kernels, which is what actually goes against the spirit of Kaggle.</p>\n\n<p>But why would a bunch of people moving up the leaderboards by running a kernel bother you?</p>\n\n<p>If someone is here purely for the competition, they shouldn't complain about other people getting better scores (regardless of how they got them). Compete harder!</p>\n\n<p>If someone is not here purely for the competition but also to learn things, well... here's a kernel that you can learn from! Does the existence of this kernel mean that you just did a bunch of work for nothing while other people get a free lunch? I don't think so. You can take the useful bits from this kernel and use them to improve your own solution.</p>\n\n<p>(For what it's worth, I do think it's a jerk move to publish a very high-scoring kernel in the last days of the competition. But again, this kernel isn't really that great -- logloss of 0.46, come on -- and we're only halfway.)</p>",
      "rawMarkdown": "To me, the spirit of Kaggle is about learning to get better at doing data science together.\n\nPeople who just run high-scoring kernels to get a position on the leaderboard, is a side-effect of allowing people to share kernels. It may not be ideal, but it's also not a big deal. A couple of weeks ago, a lot of people shot up on the leaderboards because they all ran the script that exploited the leakage. It happens. The only way to avoid that is to disable public kernels, which is what actually goes against the spirit of Kaggle.\n\nBut why would a bunch of people moving up the leaderboards by running a kernel bother you?\n\nIf someone is here purely for the competition, they shouldn't complain about other people getting better scores (regardless of how they got them). Compete harder!\n\nIf someone is not here purely for the competition but also to learn things, well... here's a kernel that you can learn from! Does the existence of this kernel mean that you just did a bunch of work for nothing while other people get a free lunch? I don't think so. You can take the useful bits from this kernel and use them to improve your own solution.\n\n(For what it's worth, I do think it's a jerk move to publish a very high-scoring kernel in the last days of the competition. But again, this kernel isn't really that great -- logloss of 0.46, come on -- and we're only halfway.)",
      "votes": null
    },
    {
      "id": "729686",
      "postDate": "01/26/2020 14:22:50",
      "content": "<p>Oh, and also: here's the corresponding <a href=\"https://www.kaggle.com/humananalog/binary-image-classifier-training-demo\">training kernel</a>. ;-)</p>",
      "rawMarkdown": "Oh, and also: here's the corresponding [training kernel](https://www.kaggle.com/humananalog/binary-image-classifier-training-demo). ;-)",
      "votes": null
    },
    {
      "id": "729695",
      "postDate": "01/26/2020 14:33:36",
      "content": "<p>I just genuinely believe people learn much less from a copy-and-paste kernel they submit, rather than a few tips which forces them to adapt themselves. I mean one doesn't need to give them the models, for example, for them to learn how to write good prediction code. If they were taking your prediction code and adapting it to their own models I would wager that would be more educational. But it's fine, we'll just have to agree to disagree.</p>\n\n<p>You're right with 2 months left it probably doesn't matter, but it is likely to reduce heterogeneity in the population rather than increase. In the BengaliAI competition there is a bronze kernel published which has led to over 20% of the competition have the same score.</p>",
      "rawMarkdown": "I just genuinely believe people learn much less from a copy-and-paste kernel they submit, rather than a few tips which forces them to adapt themselves. I mean one doesn't need to give them the models, for example, for them to learn how to write good prediction code. If they were taking your prediction code and adapting it to their own models I would wager that would be more educational. But it's fine, we'll just have to agree to disagree.\n\nYou're right with 2 months left it probably doesn't matter, but it is likely to reduce heterogeneity in the population rather than increase. In the BengaliAI competition there is a bronze kernel published which has led to over 20% of the competition have the same score.",
      "votes": null
    },
    {
      "id": "729753",
      "postDate": "01/26/2020 16:01:20",
      "content": "<blockquote>\n  <p>I just genuinely believe people learn much less from a copy-and-paste kernel they submit, rather than a few tips which forces them to adapt themselves</p>\n</blockquote>\n\n<p>In general I'd say I agree with that. But this competition is harder to get started with than most (due to the large volume of data). Should it only be accessible to people with good coding skills? </p>\n\n<p>Or should we level the playing field by giving people a working example that they can use as a baseline, and hereby include people who would normally not have a chance because the barrier of entry was too high for them, or who would have given up because their submissions kept failing for mysterious reasons?</p>\n\n<p>Including the trained checkpoint was a judgment call. Yes, it means a bunch of people will now have a meaningless score that is cluttering up the leaderboard (for a short while). But it also proves that the kernel works. If I had included an untrained checkpoint that scores 0.6931 or worse, is that because of a bug in the kernel or because the model is just bad?</p>\n\n<p>Anyway, not writing this to argue endlessly, just to give an insight into why I decided to publish this kernel.</p>",
      "rawMarkdown": "&gt; I just genuinely believe people learn much less from a copy-and-paste kernel they submit, rather than a few tips which forces them to adapt themselves\n\nIn general I'd say I agree with that. But this competition is harder to get started with than most (due to the large volume of data). Should it only be accessible to people with good coding skills? \n\nOr should we level the playing field by giving people a working example that they can use as a baseline, and hereby include people who would normally not have a chance because the barrier of entry was too high for them, or who would have given up because their submissions kept failing for mysterious reasons?\n\nIncluding the trained checkpoint was a judgment call. Yes, it means a bunch of people will now have a meaningless score that is cluttering up the leaderboard (for a short while). But it also proves that the kernel works. If I had included an untrained checkpoint that scores 0.6931 or worse, is that because of a bug in the kernel or because the model is just bad?\n\nAnyway, not writing this to argue endlessly, just to give an insight into why I decided to publish this kernel.",
      "votes": null
    },
    {
      "id": "729755",
      "postDate": "01/26/2020 16:04:01",
      "content": "<p>All fair points! And congratulations on your kernel, too!</p>",
      "rawMarkdown": "All fair points! And congratulations on your kernel, too!",
      "votes": null
    },
    {
      "id": "729887",
      "postDate": "01/26/2020 19:11:17",
      "content": "<p>Well, to be honest, I couldn't care less if people copy pasted some kernel and got their name on the top 50 LB. But this comment is to primarily reply to your question - \"Should it only be accessible to people with good coding skills?\"</p>\n\n<p>In my opinion, not having a solution forces people to improve their coding skills and therefore take on challenges which they otherwise wouldn't. I personally am not the best coder around. However, because of this competition, I have pretty much worked for long hours making my own pipeline from scratch and it feels great. Sure, I haven't even submitted a kernel yet, or even made an attempt to. But even without doing so, I have made so much personal progress on my skill level.</p>",
      "rawMarkdown": "Well, to be honest, I couldn't care less if people copy pasted some kernel and got their name on the top 50 LB. But this comment is to primarily reply to your question - \"Should it only be accessible to people with good coding skills?\"\n\nIn my opinion, not having a solution forces people to improve their coding skills and therefore take on challenges which they otherwise wouldn't. I personally am not the best coder around. However, because of this competition, I have pretty much worked for long hours making my own pipeline from scratch and it feels great. Sure, I haven't even submitted a kernel yet, or even made an attempt to. But even without doing so, I have made so much personal progress on my skill level.",
      "votes": null
    },
    {
      "id": "729898",
      "postDate": "01/26/2020 19:26:01",
      "content": "<pre><code>...posted earlier in wrong thread\nMy main objection to the Kernel was that it was dependent on a training checkpoint that was provided by the author without the code that produced it. Surely a more poorly performing model could have been supplied to preserve the LB. It seemed to me that this guaranteed a large cohort that would run it and not improve on it, much like a leak (as has been mentioned). I thought that was a bad idea. But now that HA (with significant effort) has provided the training kernel, that objection is no longer there so I'm now fine with the kernel. I have myself benefited in the past from good kernels that I went on to improve.\nI won't run it (because I like my own code, poorly performing as it is) but I will read it carefully and use the good stuff.\n</code></pre>",
      "rawMarkdown": "...posted earlier in wrong thread\n    My main objection to the Kernel was that it was dependent on a training checkpoint that was provided by the author without the code that produced it. Surely a more poorly performing model could have been supplied to preserve the LB. It seemed to me that this guaranteed a large cohort that would run it and not improve on it, much like a leak (as has been mentioned). I thought that was a bad idea. But now that HA (with significant effort) has provided the training kernel, that objection is no longer there so I'm now fine with the kernel. I have myself benefited in the past from good kernels that I went on to improve.\n    I won't run it (because I like my own code, poorly performing as it is) but I will read it carefully and use the good stuff.",
      "votes": null
    },
    {
      "id": "730019",
      "postDate": "01/27/2020 01:35:34",
      "content": "<p>Oversharing is a recurrent problem on Kaggle...</p>",
      "rawMarkdown": "Oversharing is a recurrent problem on Kaggle...",
      "votes": null
    },
    {
      "id": "730489",
      "postDate": "01/27/2020 14:48:43",
      "content": "<p>I want to frame my words here as respectfully as I can, I am pretty sure <a href=\"/humananalog\">@humananalog</a> did this in good spirit, however also want to voice my disagreement to the sharing the original model.</p>\n\n<p>in one hand I am very much against sharing high scoring kernel - let's, for now, define \"high-scoring\" as top 5% LB - when a competition is well underway. this one has been ongoing for more than 6 weeks, for me losing of the  LB at this point as a reflective mechanism take too much of the fun out of the competition - it is not good for kaggle as a platform if we encounter this almost every competition. we also should take into consideration that this competition already survived a leak</p>\n\n<p>additionally, I especially dislike sharing of high scoring pre-train model when there is no way to reproduce. I understand that <a href=\"/humananalog\">@humananalog</a> has already shared a training kernel with a simplified version, I still take the view that few would be able to replicate his shared model. I agree with <a href=\"/feifeizaici\">@feifeizaici</a> that this is too high a score for a baseline model at the current point. Time will tell, I am prepared to be proven wrong on this one. </p>\n\n<p>In the other hand, I really appreciate <a href=\"/humananalog\">@humananalog</a> effort in making the inference pipeline available, and the subsequent effort publishes part of his training scheme - nice touches, and I personally am learning from his work. </p>\n\n<p>so that is it, I would stop bitching about it in this competition forum, and suffer more to improve my own modelling pipeline. </p>",
      "rawMarkdown": "I want to frame my words here as respectfully as I can, I am pretty sure @humananalog did this in good spirit, however also want to voice my disagreement to the sharing the original model.\n\nin one hand I am very much against sharing high scoring kernel - let's, for now, define \"high-scoring\" as top 5% LB - when a competition is well underway. this one has been ongoing for more than 6 weeks, for me losing of the  LB at this point as a reflective mechanism take too much of the fun out of the competition - it is not good for kaggle as a platform if we encounter this almost every competition. we also should take into consideration that this competition already survived a leak\n\nadditionally, I especially dislike sharing of high scoring pre-train model when there is no way to reproduce. I understand that @humananalog has already shared a training kernel with a simplified version, I still take the view that few would be able to replicate his shared model. I agree with @feifeizaici that this is too high a score for a baseline model at the current point. Time will tell, I am prepared to be proven wrong on this one. \n\nIn the other hand, I really appreciate @humananalog effort in making the inference pipeline available, and the subsequent effort publishes part of his training scheme - nice touches, and I personally am learning from his work. \n\nso that is it, I would stop bitching about it in this competition forum, and suffer more to improve my own modelling pipeline.",
      "votes": null
    },
    {
      "id": "730509",
      "postDate": "01/27/2020 15:08:43",
      "content": "<p>Thank for sharing bro, I'm just a beginner and I learn a lot from this site. I'm just coming up with simple ideas to implement, sometime people come up with complex kernels and you don't learn as much if you just go through and implement a simple model.</p>\n\n<p>Will go through your submissions today. Thanks for this thread, I wouldn't have noticed it otherwise</p>",
      "rawMarkdown": "Thank for sharing bro, I'm just a beginner and I learn a lot from this site. I'm just coming up with simple ideas to implement, sometime people come up with complex kernels and you don't learn as much if you just go through and implement a simple model.\n\nWill go through your submissions today. Thanks for this thread, I wouldn't have noticed it otherwise",
      "votes": null
    },
    {
      "id": "731313",
      "postDate": "01/28/2020 14:01:42",
      "content": "<p>Agree, he took the fun in the game.</p>",
      "rawMarkdown": "Agree, he took the fun in the game.",
      "votes": null
    },
    {
      "id": "731428",
      "postDate": "01/28/2020 16:22:34",
      "content": "<p>Yet, only three days after I posted the inference kernel, the people who used it are quickly dropping off the first page of the leaderboard... So it doesn't seem to have done too much harm (and perhaps has even encouraged people to improve their solutions, who knows).</p>",
      "rawMarkdown": "Yet, only three days after I posted the inference kernel, the people who used it are quickly dropping off the first page of the leaderboard... So it doesn't seem to have done too much harm (and perhaps has even encouraged people to improve their solutions, who knows).",
      "votes": null
    },
    {
      "id": "731473",
      "postDate": "01/28/2020 17:15:07",
      "content": "<p>Not sure I agree with your conclusion.</p>\n\n<p>My conclusion would be that folks found ways to improve your kernel.   Looking at the top 100 my bet would be those folks with less than 5 submissions are mostly using your kernel.  </p>",
      "rawMarkdown": "Not sure I agree with your conclusion.\n\nMy conclusion would be that folks found ways to improve your kernel.   Looking at the top 100 my bet would be those folks with less than 5 submissions are mostly using your kernel.",
      "votes": null
    },
    {
      "id": "731479",
      "postDate": "01/28/2020 17:22:50",
      "content": "<p>To put things in perspective a little, 96 people currently have your kernel's score. PC Jimmy here, without them, would be in 59th place rather than 155th, if we assume their scores would've otherwise been lower than his.</p>\n\n<p>This corresponds with him going from silver medal to no medal.</p>",
      "rawMarkdown": "To put things in perspective a little, 96 people currently have your kernel's score. PC Jimmy here, without them, would be in 59th place rather than 155th, if we assume their scores would've otherwise been lower than his.\n\nThis corresponds with him going from silver medal to no medal.",
      "votes": null
    },
    {
      "id": "731489",
      "postDate": "01/28/2020 17:35:03",
      "content": "<p><a href=\"/humananalog\">@humananalog</a> in the spirit of sharing. Here are some frames. <a href=\"https://www.kaggle.com/sciarrilli/dfdc-f150\">https://www.kaggle.com/sciarrilli/dfdc-f150</a> \nthanks for the notebooks by the way. great for learning. </p>",
      "rawMarkdown": "humananalog in the spirit of sharing. Here are some frames. https://www.kaggle.com/sciarrilli/dfdc-f150 \nthanks for the notebooks by the way. great for learning.",
      "votes": null
    },
    {
      "id": "731533",
      "postDate": "01/28/2020 19:04:22",
      "content": "<p>Valid point, if it were the last day of the competition. Three days ago, running my kernel got you to position 30. Now it's position 75. By the end of the week, I expect those people to be out of the silver medals too.</p>",
      "rawMarkdown": "Valid point, if it were the last day of the competition. Three days ago, running my kernel got you to position 30. Now it's position 75. By the end of the week, I expect those people to be out of the silver medals too.",
      "votes": null
    },
    {
      "id": "731535",
      "postDate": "01/28/2020 19:05:39",
      "content": "<p><a href=\"/humananalog\">@humananalog</a> That's the case for me. I had a look to your model and tried it because my model (with similar approach) did not work so well for me (LB 0.51). Then, it showed me that my model was working fine indeed with some little adjustments (LB 0.45). One difficulty is to understand why such model overfits so quickly. We're far from the end of the competition, IMO your kernel gave a boost to force everyone to improve.</p>",
      "rawMarkdown": "humananalog That's the case for me. I had a look to your model and tried it because my model (with similar approach) did not work so well for me (LB 0.51). Then, it showed me that my model was working fine indeed with some little adjustments (LB 0.45). One difficulty is to understand why such model overfits so quickly. We're far from the end of the competition, IMO your kernel gave a boost to force everyone to improve.",
      "votes": null
    },
    {
      "id": "731620",
      "postDate": "01/28/2020 20:22:51",
      "content": "<p>Sharing leads to more/faster progress, and it likely means a better overall logloss when the final models are submitted. I don't see it as a problem.</p>",
      "rawMarkdown": "Sharing leads to more/faster progress, and it likely means a better overall logloss when the final models are submitted. I don't see it as a problem.",
      "votes": null
    },
    {
      "id": "731630",
      "postDate": "01/28/2020 20:53:18",
      "content": "<p>James Howard - actually if Human Analog had not shared his kernel I probably would be in IDGAF land working on a different competition.   After the first month of just getting a submission to finally work I was ready to call it quits as I realized the difficulties I was facing getting ffmpeg to work the way I wanted.  His kernel showed me a different set of tools that mesh nicely with the things I wanted to do using ffmpeg/mtcnn.  I am back to having hope that I can learn and generate new code every day.</p>\n\n<p>As I type this we are sitting on 64 days left - all the folks who are throwing shade about his share - I don't recall a single competition I have worked on over the last year were we would have been sitting at starter kits for as long as this one did.  For  most of that past it seemed like at around the 30 days left to go;  was the point when shares started drying up.  This competition seemed to have dried up in the first week until his share.  </p>\n\n<p>So with his share I am sitting at 155 and without his share I would be sitting on the couch watching the Impeachment of President Trump.  He saved from flipping between CNN AND Fox News trying to compute the average.</p>",
      "rawMarkdown": "James Howard - actually if Human Analog had not shared his kernel I probably would be in IDGAF land working on a different competition.   After the first month of just getting a submission to finally work I was ready to call it quits as I realized the difficulties I was facing getting ffmpeg to work the way I wanted.  His kernel showed me a different set of tools that mesh nicely with the things I wanted to do using ffmpeg/mtcnn.  I am back to having hope that I can learn and generate new code every day.\n\nAs I type this we are sitting on 64 days left - all the folks who are throwing shade about his share - I don't recall a single competition I have worked on over the last year were we would have been sitting at starter kits for as long as this one did.  For  most of that past it seemed like at around the 30 days left to go;  was the point when shares started drying up.  This competition seemed to have dried up in the first week until his share.  \n\nSo with his share I am sitting at 155 and without his share I would be sitting on the couch watching the Impeachment of President Trump.  He saved from flipping between CNN AND Fox News trying to compute the average.",
      "votes": null
    },
    {
      "id": "731726",
      "postDate": "01/29/2020 01:04:16",
      "content": "<p>Sharing is not a problem. Oversharing is.</p>",
      "rawMarkdown": "Sharing is not a problem. Oversharing is.",
      "votes": null
    },
    {
      "id": "731758",
      "postDate": "01/29/2020 02:26:56",
      "content": "<p>I'm very thankful to Human Analog for sharing this kernel. I've been meaning to learn Pytorch for months now, and his kernel gave me the opportunity to re-open that pytorch book, and learn everything I've been delaying. I'm not here to win competitions, I'm here to learn. Thanks for sharing ! </p>",
      "rawMarkdown": "I'm very thankful to Human Analog for sharing this kernel. I've been meaning to learn Pytorch for months now, and his kernel gave me the opportunity to re-open that pytorch book, and learn everything I've been delaying. I'm not here to win competitions, I'm here to learn. Thanks for sharing !",
      "votes": null
    },
    {
      "id": "733111",
      "postDate": "01/30/2020 17:12:07",
      "content": "<p>Human Analog had a goodwill to share.\nBut,  the process of deep learning is veiled and somewhat ambiguous yet.</p>\n\n<p>I experienced that some models work well, on the other hand other models didn't.\nAnd I can't explain it clearly.\nSo, I think kaggle competitions include all that painful labor.</p>\n\n<p>There exists grey-zone between encourage to compete and learn.\nBut I think sharing the pre-trained weights scoring 100th in LB was too far \nalthough it left more than a month to the end of DFDC.</p>\n\n<p>It's not pleasant to see just-copy-and-shiftenter player beat participants who\nhave been made big effort over than a month.</p>\n\n<p>I had been worse than 0.46788 for a long time and many participants on the way to\nimprove their model but now just stall on that score.</p>\n\n<p>Rank is Rank.\nIt's more than a copy and paste.</p>",
      "rawMarkdown": "Human Analog had a goodwill to share.\nBut,  the process of deep learning is veiled and somewhat ambiguous yet.\n\nI experienced that some models work well, on the other hand other models didn't.\nAnd I can't explain it clearly.\nSo, I think kaggle competitions include all that painful labor.\n\nThere exists grey-zone between encourage to compete and learn.\nBut I think sharing the pre-trained weights scoring 100th in LB was too far \nalthough it left more than a month to the end of DFDC.\n\nIt's not pleasant to see just-copy-and-shiftenter player beat participants who\nhave been made big effort over than a month.\n\nI had been worse than 0.46788 for a long time and many participants on the way to\nimprove their model but now just stall on that score.\n\nRank is Rank.\nIt's more than a copy and paste.",
      "votes": null
    },
    {
      "id": "733326",
      "postDate": "01/31/2020 01:07:55",
      "content": "<p><a href=\"/humananalog\">@humananalog</a> Thank you for your generous contributions to this competition. It's interesting to learn that the model in the inference demo was only trained using ~30000 images. Any chance you can tell us if some of these images were sampled from compressed videos?</p>",
      "rawMarkdown": "humananalog Thank you for your generous contributions to this competition. It's interesting to learn that the model in the inference demo was only trained using ~30000 images. Any chance you can tell us if some of these images were sampled from compressed videos?",
      "votes": null
    },
    {
      "id": "733643",
      "postDate": "01/31/2020 11:34:44",
      "content": "<p>I did say ~30000 images but I just looked at the code for that particular model and it's not entirely correct. This was for an older version of the model that performed in the 0.55xxx range. </p>\n\n<p>The training set I used for the 0.46xxx model has 16815 real videos and 85873 fake videos. In each batch, I sample half the batch from the ~17K real videos. The other half is sampled from the ~86K fake videos. So this actually uses ~100K images.</p>\n\n<p>Part of the image augmentation I do involves randomly compressing the image as JPEG with a randomly chosen quality setting. This approximates video compression.</p>",
      "rawMarkdown": "I did say ~30000 images but I just looked at the code for that particular model and it's not entirely correct. This was for an older version of the model that performed in the 0.55xxx range. \n\nThe training set I used for the 0.46xxx model has 16815 real videos and 85873 fake videos. In each batch, I sample half the batch from the ~17K real videos. The other half is sampled from the ~86K fake videos. So this actually uses ~100K images.\n\nPart of the image augmentation I do involves randomly compressing the image as JPEG with a randomly chosen quality setting. This approximates video compression.",
      "votes": null
    },
    {
      "id": "733677",
      "postDate": "01/31/2020 12:14:04",
      "content": "<blockquote>\n  <p><strong>Human Analog wrote:</strong>\n  The training set I used for the 0.46xxx model has 16815 real videos and 85873 fake videos. In each batch, I sample half the batch from the ~17K real videos. The other half is sampled from the ~86K fake videos. So this actually uses ~100K images.</p>\n  \n  <p>Part of the image augmentation I do involves randomly compressing the image as JPEG with a randomly chosen quality setting. This approximates video compression.</p>\n</blockquote>\n\n<p>Thanks for making this clarification, on my side with my own dataset, my observation is that I can get to about 0.51LB with 50k images, and probably to a similar level to 0.45LB with 100k - I haven't check for models using this specific amount of training data. </p>\n\n<p>It is reassuring to see other's LB score and training data amount is somehow correlate to our own. </p>",
      "rawMarkdown": "&gt; **Human Analog wrote:**\n&gt; The training set I used for the 0.46xxx model has 16815 real videos and 85873 fake videos. In each batch, I sample half the batch from the ~17K real videos. The other half is sampled from the ~86K fake videos. So this actually uses ~100K images.\n&gt; \n&gt; Part of the image augmentation I do involves randomly compressing the image as JPEG with a randomly chosen quality setting. This approximates video compression.\n\nThanks for making this clarification, on my side with my own dataset, my observation is that I can get to about 0.51LB with 50k images, and probably to a similar level to 0.45LB with 100k - I haven't check for models using this specific amount of training data. \n\nIt is reassuring to see other's LB score and training data amount is somehow correlate to our own.",
      "votes": null
    },
    {
      "id": "733692",
      "postDate": "01/31/2020 12:41:16",
      "content": "<p>My biggest gain came from combining the real and fake images inside the same batch, so that the model always sees a balanced amount. (This doesn't choose the fake images randomly, but samples from the fakes that belong to the reals that have been chosen from the batch.)</p>",
      "rawMarkdown": "My biggest gain came from combining the real and fake images inside the same batch, so that the model always sees a balanced amount. (This doesn't choose the fake images randomly, but samples from the fakes that belong to the reals that have been chosen from the batch.)",
      "votes": null
    },
    {
      "id": "733695",
      "postDate": "01/31/2020 12:44:38",
      "content": "<p>interesting, and thanks for sharing. \nfor me the largest gain so far come from 1) increasing amount of training data to a certain level, and 2) getting augmentation right - i.e. not too heavy and not too light</p>\n\n<p>and of course, having a reliable inference pipleline is also critical - I am sure we agree on this one 😉 </p>",
      "rawMarkdown": "interesting, and thanks for sharing. \nfor me the largest gain so far come from 1) increasing amount of training data to a certain level, and 2) getting augmentation right - i.e. not too heavy and not too light\n\nand of course, having a reliable inference pipleline is also critical - I am sure we agree on this one 😉",
      "votes": null
    },
    {
      "id": "734286",
      "postDate": "02/01/2020 08:05:15",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "734518",
      "postDate": "02/01/2020 15:51:51",
      "content": "<p>sometimes I train the whole network, sometime I do some fine-tune before training the whole network - my ad-hoc observation seems to suggest that this works better with smaller training size, \naside from that, it really depends on my mood.... :)</p>",
      "rawMarkdown": "sometimes I train the whole network, sometime I do some fine-tune before training the whole network - my ad-hoc observation seems to suggest that this works better with smaller training size, \naside from that, it really depends on my mood.... :)",
      "votes": null
    },
    {
      "id": "734776",
      "postDate": "02/02/2020 01:47:40",
      "content": "<p><a href=\"/humananalog\">@humananalog</a> Awesome, thanks for the advice</p>",
      "rawMarkdown": "humananalog Awesome, thanks for the advice",
      "votes": null
    },
    {
      "id": "748382",
      "postDate": "02/17/2020 13:19:03",
      "content": "<p>That's what happens when sharing is gamified :)</p>",
      "rawMarkdown": "That's what happens when sharing is gamified :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 729224,
      "author_name": "naivelamb",
      "author_url": "",
      "post_date": "01/25/2020 23:17:51",
      "content": "<p>Check <a href=\"https://www.kaggle.com/humananalog/inference-demo\">here</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 729225,
          "author_name": "petewills",
          "author_url": "",
          "post_date": "01/25/2020 23:19:11",
          "content": "<p>Yes, just saw it thanks. I missed it because it wasn't flagged on the LB. Bummer to have to compete with a silver kernel.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 729238,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "01/25/2020 23:50:09",
          "content": "<p>I'm the author of that kernel. 0.46788 really isn't that great of score on a logistic loss metric, so don't feel too discouraged...</p>\n\n<p>I suggest using this kernel as a baseline. If you use the same process with your own model (i.e. extract 17 or so frames from each video, do a prediction on each frame, take the mean) but your model scores less than 0.46x, it's a sign you need to improve your training data. </p>\n\n<p>As I said in the description, this uses a very simple ResNeXt50 model. It was only trained for about 7 epochs on ~30000 training images (split evenly between real/fake). I didn't spend any time on getting the best possible model design or training procedure. Each of my improvements on the leaderboard was done by improving my set of training images.</p>\n\n<p>So if you take the same approach as this kernel and your score is worse, spend some more time improving your dataset. Keep doing that and I guarantee you'll get a score below 0.46x. 😄 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 729239,
          "author_name": "naivelamb",
          "author_url": "",
          "post_date": "01/25/2020 23:51:33",
          "content": "<p>Considering there are 2 months left, I don't think 0.46788 would give any medals at the end. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 729257,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "01/26/2020 00:31:52",
          "content": "<p>Been in a lot of competitions over the past two years - I don't recall that I have ever seen a public shared kernel - at the two month point - be in the metal.</p>\n\n<p>I was very worried that the prize money would keep folks from posting good kernels.  Until Human Analog posted his, all the shared kernels had a leader board score that was pretty much the same that you can get with a random number generated submission.</p>\n\n<p>Hoping that his share will kick a few more loose rocks out of the mountain - pretty hard for the community to help in the creation of a worthy product when we are keeping our cards hidden.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 733326,
          "author_name": "jamesthornton",
          "author_url": "",
          "post_date": "01/31/2020 01:07:55",
          "content": "<p><a href=\"/humananalog\">@humananalog</a> Thank you for your generous contributions to this competition. It's interesting to learn that the model in the inference demo was only trained using ~30000 images. Any chance you can tell us if some of these images were sampled from compressed videos?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 733643,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "01/31/2020 11:34:44",
          "content": "<p>I did say ~30000 images but I just looked at the code for that particular model and it's not entirely correct. This was for an older version of the model that performed in the 0.55xxx range. </p>\n\n<p>The training set I used for the 0.46xxx model has 16815 real videos and 85873 fake videos. In each batch, I sample half the batch from the ~17K real videos. The other half is sampled from the ~86K fake videos. So this actually uses ~100K images.</p>\n\n<p>Part of the image augmentation I do involves randomly compressing the image as JPEG with a randomly chosen quality setting. This approximates video compression.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 733677,
          "author_name": "yifanxie",
          "author_url": "",
          "post_date": "01/31/2020 12:14:04",
          "content": "<blockquote>\n  <p><strong>Human Analog wrote:</strong>\n  The training set I used for the 0.46xxx model has 16815 real videos and 85873 fake videos. In each batch, I sample half the batch from the ~17K real videos. The other half is sampled from the ~86K fake videos. So this actually uses ~100K images.</p>\n  \n  <p>Part of the image augmentation I do involves randomly compressing the image as JPEG with a randomly chosen quality setting. This approximates video compression.</p>\n</blockquote>\n\n<p>Thanks for making this clarification, on my side with my own dataset, my observation is that I can get to about 0.51LB with 50k images, and probably to a similar level to 0.45LB with 100k - I haven't check for models using this specific amount of training data. </p>\n\n<p>It is reassuring to see other's LB score and training data amount is somehow correlate to our own. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 733692,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "01/31/2020 12:41:16",
          "content": "<p>My biggest gain came from combining the real and fake images inside the same batch, so that the model always sees a balanced amount. (This doesn't choose the fake images randomly, but samples from the fakes that belong to the reals that have been chosen from the batch.)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 733695,
          "author_name": "yifanxie",
          "author_url": "",
          "post_date": "01/31/2020 12:44:38",
          "content": "<p>interesting, and thanks for sharing. \nfor me the largest gain so far come from 1) increasing amount of training data to a certain level, and 2) getting augmentation right - i.e. not too heavy and not too light</p>\n\n<p>and of course, having a reliable inference pipleline is also critical - I am sure we agree on this one 😉 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 734286,
          "author_name": "cslaker",
          "author_url": "",
          "post_date": "02/01/2020 08:05:15",
          "content": "",
          "votes": null,
          "replies": []
        },
        {
          "id": 734518,
          "author_name": "yifanxie",
          "author_url": "",
          "post_date": "02/01/2020 15:51:51",
          "content": "<p>sometimes I train the whole network, sometime I do some fine-tune before training the whole network - my ad-hoc observation seems to suggest that this works better with smaller training size, \naside from that, it really depends on my mood.... :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 734776,
          "author_name": "jamesthornton",
          "author_url": "",
          "post_date": "02/02/2020 01:47:40",
          "content": "<p><a href=\"/humananalog\">@humananalog</a> Awesome, thanks for the advice</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 729347,
      "author_name": "feifeizaici",
      "author_url": "",
      "post_date": "01/26/2020 04:34:30",
      "content": "<p>Actually I don't think 0.46788 is a good baseline. It's too high in LB.\nPS. Tips and a good pipeline(like yours) are already enough for starters. But a trained model made the LB useless and it reminded me of the leak.</p>",
      "votes": null,
      "replies": [
        {
          "id": 729595,
          "author_name": "jamesphoward",
          "author_url": "",
          "post_date": "01/26/2020 12:31:09",
          "content": "<p>Agreed, people who copy this kernel immediately beat 90% of entrants who will have done an order of magnitude more work. It's against the spirit of Kaggle.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 729684,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "01/26/2020 14:21:15",
          "content": "<p>To me, the spirit of Kaggle is about learning to get better at doing data science together.</p>\n\n<p>People who just run high-scoring kernels to get a position on the leaderboard, is a side-effect of allowing people to share kernels. It may not be ideal, but it's also not a big deal. A couple of weeks ago, a lot of people shot up on the leaderboards because they all ran the script that exploited the leakage. It happens. The only way to avoid that is to disable public kernels, which is what actually goes against the spirit of Kaggle.</p>\n\n<p>But why would a bunch of people moving up the leaderboards by running a kernel bother you?</p>\n\n<p>If someone is here purely for the competition, they shouldn't complain about other people getting better scores (regardless of how they got them). Compete harder!</p>\n\n<p>If someone is not here purely for the competition but also to learn things, well... here's a kernel that you can learn from! Does the existence of this kernel mean that you just did a bunch of work for nothing while other people get a free lunch? I don't think so. You can take the useful bits from this kernel and use them to improve your own solution.</p>\n\n<p>(For what it's worth, I do think it's a jerk move to publish a very high-scoring kernel in the last days of the competition. But again, this kernel isn't really that great -- logloss of 0.46, come on -- and we're only halfway.)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 729686,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "01/26/2020 14:22:50",
          "content": "<p>Oh, and also: here's the corresponding <a href=\"https://www.kaggle.com/humananalog/binary-image-classifier-training-demo\">training kernel</a>. ;-)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 729695,
          "author_name": "jamesphoward",
          "author_url": "",
          "post_date": "01/26/2020 14:33:36",
          "content": "<p>I just genuinely believe people learn much less from a copy-and-paste kernel they submit, rather than a few tips which forces them to adapt themselves. I mean one doesn't need to give them the models, for example, for them to learn how to write good prediction code. If they were taking your prediction code and adapting it to their own models I would wager that would be more educational. But it's fine, we'll just have to agree to disagree.</p>\n\n<p>You're right with 2 months left it probably doesn't matter, but it is likely to reduce heterogeneity in the population rather than increase. In the BengaliAI competition there is a bronze kernel published which has led to over 20% of the competition have the same score.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 729753,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "01/26/2020 16:01:20",
          "content": "<blockquote>\n  <p>I just genuinely believe people learn much less from a copy-and-paste kernel they submit, rather than a few tips which forces them to adapt themselves</p>\n</blockquote>\n\n<p>In general I'd say I agree with that. But this competition is harder to get started with than most (due to the large volume of data). Should it only be accessible to people with good coding skills? </p>\n\n<p>Or should we level the playing field by giving people a working example that they can use as a baseline, and hereby include people who would normally not have a chance because the barrier of entry was too high for them, or who would have given up because their submissions kept failing for mysterious reasons?</p>\n\n<p>Including the trained checkpoint was a judgment call. Yes, it means a bunch of people will now have a meaningless score that is cluttering up the leaderboard (for a short while). But it also proves that the kernel works. If I had included an untrained checkpoint that scores 0.6931 or worse, is that because of a bug in the kernel or because the model is just bad?</p>\n\n<p>Anyway, not writing this to argue endlessly, just to give an insight into why I decided to publish this kernel.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 729755,
          "author_name": "jamesphoward",
          "author_url": "",
          "post_date": "01/26/2020 16:04:01",
          "content": "<p>All fair points! And congratulations on your kernel, too!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 729887,
          "author_name": "akashnandi",
          "author_url": "",
          "post_date": "01/26/2020 19:11:17",
          "content": "<p>Well, to be honest, I couldn't care less if people copy pasted some kernel and got their name on the top 50 LB. But this comment is to primarily reply to your question - \"Should it only be accessible to people with good coding skills?\"</p>\n\n<p>In my opinion, not having a solution forces people to improve their coding skills and therefore take on challenges which they otherwise wouldn't. I personally am not the best coder around. However, because of this competition, I have pretty much worked for long hours making my own pipeline from scratch and it feels great. Sure, I haven't even submitted a kernel yet, or even made an attempt to. But even without doing so, I have made so much personal progress on my skill level.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 730509,
          "author_name": "joe1dataset",
          "author_url": "",
          "post_date": "01/27/2020 15:08:43",
          "content": "<p>Thank for sharing bro, I'm just a beginner and I learn a lot from this site. I'm just coming up with simple ideas to implement, sometime people come up with complex kernels and you don't learn as much if you just go through and implement a simple model.</p>\n\n<p>Will go through your submissions today. Thanks for this thread, I wouldn't have noticed it otherwise</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 729898,
      "author_name": "petewills",
      "author_url": "",
      "post_date": "01/26/2020 19:26:01",
      "content": "<pre><code>...posted earlier in wrong thread\nMy main objection to the Kernel was that it was dependent on a training checkpoint that was provided by the author without the code that produced it. Surely a more poorly performing model could have been supplied to preserve the LB. It seemed to me that this guaranteed a large cohort that would run it and not improve on it, much like a leak (as has been mentioned). I thought that was a bad idea. But now that HA (with significant effort) has provided the training kernel, that objection is no longer there so I'm now fine with the kernel. I have myself benefited in the past from good kernels that I went on to improve.\nI won't run it (because I like my own code, poorly performing as it is) but I will read it carefully and use the good stuff.\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 730019,
      "author_name": "felipefonte99",
      "author_url": "",
      "post_date": "01/27/2020 01:35:34",
      "content": "<p>Oversharing is a recurrent problem on Kaggle...</p>",
      "votes": null,
      "replies": [
        {
          "id": 731620,
          "author_name": "meicher",
          "author_url": "",
          "post_date": "01/28/2020 20:22:51",
          "content": "<p>Sharing leads to more/faster progress, and it likely means a better overall logloss when the final models are submitted. I don't see it as a problem.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 731726,
          "author_name": "felipefonte99",
          "author_url": "",
          "post_date": "01/29/2020 01:04:16",
          "content": "<p>Sharing is not a problem. Oversharing is.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 748382,
          "author_name": "marcogorelli",
          "author_url": "",
          "post_date": "02/17/2020 13:19:03",
          "content": "<p>That's what happens when sharing is gamified :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 730489,
      "author_name": "yifanxie",
      "author_url": "",
      "post_date": "01/27/2020 14:48:43",
      "content": "<p>I want to frame my words here as respectfully as I can, I am pretty sure <a href=\"/humananalog\">@humananalog</a> did this in good spirit, however also want to voice my disagreement to the sharing the original model.</p>\n\n<p>in one hand I am very much against sharing high scoring kernel - let's, for now, define \"high-scoring\" as top 5% LB - when a competition is well underway. this one has been ongoing for more than 6 weeks, for me losing of the  LB at this point as a reflective mechanism take too much of the fun out of the competition - it is not good for kaggle as a platform if we encounter this almost every competition. we also should take into consideration that this competition already survived a leak</p>\n\n<p>additionally, I especially dislike sharing of high scoring pre-train model when there is no way to reproduce. I understand that <a href=\"/humananalog\">@humananalog</a> has already shared a training kernel with a simplified version, I still take the view that few would be able to replicate his shared model. I agree with <a href=\"/feifeizaici\">@feifeizaici</a> that this is too high a score for a baseline model at the current point. Time will tell, I am prepared to be proven wrong on this one. </p>\n\n<p>In the other hand, I really appreciate <a href=\"/humananalog\">@humananalog</a> effort in making the inference pipeline available, and the subsequent effort publishes part of his training scheme - nice touches, and I personally am learning from his work. </p>\n\n<p>so that is it, I would stop bitching about it in this competition forum, and suffer more to improve my own modelling pipeline. </p>",
      "votes": null,
      "replies": [
        {
          "id": 731313,
          "author_name": "simba29",
          "author_url": "",
          "post_date": "01/28/2020 14:01:42",
          "content": "<p>Agree, he took the fun in the game.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 731428,
      "author_name": "humananalog",
      "author_url": "",
      "post_date": "01/28/2020 16:22:34",
      "content": "<p>Yet, only three days after I posted the inference kernel, the people who used it are quickly dropping off the first page of the leaderboard... So it doesn't seem to have done too much harm (and perhaps has even encouraged people to improve their solutions, who knows).</p>",
      "votes": null,
      "replies": [
        {
          "id": 731473,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "01/28/2020 17:15:07",
          "content": "<p>Not sure I agree with your conclusion.</p>\n\n<p>My conclusion would be that folks found ways to improve your kernel.   Looking at the top 100 my bet would be those folks with less than 5 submissions are mostly using your kernel.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 731479,
          "author_name": "jamesphoward",
          "author_url": "",
          "post_date": "01/28/2020 17:22:50",
          "content": "<p>To put things in perspective a little, 96 people currently have your kernel's score. PC Jimmy here, without them, would be in 59th place rather than 155th, if we assume their scores would've otherwise been lower than his.</p>\n\n<p>This corresponds with him going from silver medal to no medal.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 731533,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "01/28/2020 19:04:22",
          "content": "<p>Valid point, if it were the last day of the competition. Three days ago, running my kernel got you to position 30. Now it's position 75. By the end of the week, I expect those people to be out of the silver medals too.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 731535,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "01/28/2020 19:05:39",
          "content": "<p><a href=\"/humananalog\">@humananalog</a> That's the case for me. I had a look to your model and tried it because my model (with similar approach) did not work so well for me (LB 0.51). Then, it showed me that my model was working fine indeed with some little adjustments (LB 0.45). One difficulty is to understand why such model overfits so quickly. We're far from the end of the competition, IMO your kernel gave a boost to force everyone to improve.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 731630,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "01/28/2020 20:53:18",
          "content": "<p>James Howard - actually if Human Analog had not shared his kernel I probably would be in IDGAF land working on a different competition.   After the first month of just getting a submission to finally work I was ready to call it quits as I realized the difficulties I was facing getting ffmpeg to work the way I wanted.  His kernel showed me a different set of tools that mesh nicely with the things I wanted to do using ffmpeg/mtcnn.  I am back to having hope that I can learn and generate new code every day.</p>\n\n<p>As I type this we are sitting on 64 days left - all the folks who are throwing shade about his share - I don't recall a single competition I have worked on over the last year were we would have been sitting at starter kits for as long as this one did.  For  most of that past it seemed like at around the 30 days left to go;  was the point when shares started drying up.  This competition seemed to have dried up in the first week until his share.  </p>\n\n<p>So with his share I am sitting at 155 and without his share I would be sitting on the couch watching the Impeachment of President Trump.  He saved from flipping between CNN AND Fox News trying to compute the average.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 731489,
      "author_name": "sciarrilli",
      "author_url": "",
      "post_date": "01/28/2020 17:35:03",
      "content": "<p><a href=\"/humananalog\">@humananalog</a> in the spirit of sharing. Here are some frames. <a href=\"https://www.kaggle.com/sciarrilli/dfdc-f150\">https://www.kaggle.com/sciarrilli/dfdc-f150</a> \nthanks for the notebooks by the way. great for learning. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 731758,
      "author_name": "jonmacpherson",
      "author_url": "",
      "post_date": "01/29/2020 02:26:56",
      "content": "<p>I'm very thankful to Human Analog for sharing this kernel. I've been meaning to learn Pytorch for months now, and his kernel gave me the opportunity to re-open that pytorch book, and learn everything I've been delaying. I'm not here to win competitions, I'm here to learn. Thanks for sharing ! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 733111,
      "author_name": "gwsong",
      "author_url": "",
      "post_date": "01/30/2020 17:12:07",
      "content": "<p>Human Analog had a goodwill to share.\nBut,  the process of deep learning is veiled and somewhat ambiguous yet.</p>\n\n<p>I experienced that some models work well, on the other hand other models didn't.\nAnd I can't explain it clearly.\nSo, I think kaggle competitions include all that painful labor.</p>\n\n<p>There exists grey-zone between encourage to compete and learn.\nBut I think sharing the pre-trained weights scoring 100th in LB was too far \nalthough it left more than a month to the end of DFDC.</p>\n\n<p>It's not pleasant to see just-copy-and-shiftenter player beat participants who\nhave been made big effort over than a month.</p>\n\n<p>I had been worse than 0.46788 for a long time and many participants on the way to\nimprove their model but now just stall on that score.</p>\n\n<p>Rank is Rank.\nIt's more than a copy and paste.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "729223": "18 submissions with exactly 0.46788 in the last 10 hours, several first time submissions. How does that happen?",
    "729224": "Check [here](https://www.kaggle.com/humananalog/inference-demo)",
    "729225": "Yes, just saw it thanks. I missed it because it wasn't flagged on the LB. Bummer to have to compete with a silver kernel.",
    "729238": "I'm the author of that kernel. 0.46788 really isn't that great of score on a logistic loss metric, so don't feel too discouraged...\n\nI suggest using this kernel as a baseline. If you use the same process with your own model (i.e. extract 17 or so frames from each video, do a prediction on each frame, take the mean) but your model scores less than 0.46x, it's a sign you need to improve your training data. \n\nAs I said in the description, this uses a very simple ResNeXt50 model. It was only trained for about 7 epochs on ~30000 training images (split evenly between real/fake). I didn't spend any time on getting the best possible model design or training procedure. Each of my improvements on the leaderboard was done by improving my set of training images.\n\nSo if you take the same approach as this kernel and your score is worse, spend some more time improving your dataset. Keep doing that and I guarantee you'll get a score below 0.46x. 😄",
    "729239": "Considering there are 2 months left, I don't think 0.46788 would give any medals at the end.",
    "729257": "Been in a lot of competitions over the past two years - I don't recall that I have ever seen a public shared kernel - at the two month point - be in the metal.\n\nI was very worried that the prize money would keep folks from posting good kernels.  Until Human Analog posted his, all the shared kernels had a leader board score that was pretty much the same that you can get with a random number generated submission.\n\nHoping that his share will kick a few more loose rocks out of the mountain - pretty hard for the community to help in the creation of a worthy product when we are keeping our cards hidden.",
    "729347": "Actually I don't think 0.46788 is a good baseline. It's too high in LB.\nPS. Tips and a good pipeline(like yours) are already enough for starters. But a trained model made the LB useless and it reminded me of the leak.",
    "729595": "Agreed, people who copy this kernel immediately beat 90% of entrants who will have done an order of magnitude more work. It's against the spirit of Kaggle.",
    "729684": "To me, the spirit of Kaggle is about learning to get better at doing data science together.\n\nPeople who just run high-scoring kernels to get a position on the leaderboard, is a side-effect of allowing people to share kernels. It may not be ideal, but it's also not a big deal. A couple of weeks ago, a lot of people shot up on the leaderboards because they all ran the script that exploited the leakage. It happens. The only way to avoid that is to disable public kernels, which is what actually goes against the spirit of Kaggle.\n\nBut why would a bunch of people moving up the leaderboards by running a kernel bother you?\n\nIf someone is here purely for the competition, they shouldn't complain about other people getting better scores (regardless of how they got them). Compete harder!\n\nIf someone is not here purely for the competition but also to learn things, well... here's a kernel that you can learn from! Does the existence of this kernel mean that you just did a bunch of work for nothing while other people get a free lunch? I don't think so. You can take the useful bits from this kernel and use them to improve your own solution.\n\n(For what it's worth, I do think it's a jerk move to publish a very high-scoring kernel in the last days of the competition. But again, this kernel isn't really that great -- logloss of 0.46, come on -- and we're only halfway.)",
    "729686": "Oh, and also: here's the corresponding [training kernel](https://www.kaggle.com/humananalog/binary-image-classifier-training-demo). ;-)",
    "729695": "I just genuinely believe people learn much less from a copy-and-paste kernel they submit, rather than a few tips which forces them to adapt themselves. I mean one doesn't need to give them the models, for example, for them to learn how to write good prediction code. If they were taking your prediction code and adapting it to their own models I would wager that would be more educational. But it's fine, we'll just have to agree to disagree.\n\nYou're right with 2 months left it probably doesn't matter, but it is likely to reduce heterogeneity in the population rather than increase. In the BengaliAI competition there is a bronze kernel published which has led to over 20% of the competition have the same score.",
    "729753": "&gt; I just genuinely believe people learn much less from a copy-and-paste kernel they submit, rather than a few tips which forces them to adapt themselves\n\nIn general I'd say I agree with that. But this competition is harder to get started with than most (due to the large volume of data). Should it only be accessible to people with good coding skills? \n\nOr should we level the playing field by giving people a working example that they can use as a baseline, and hereby include people who would normally not have a chance because the barrier of entry was too high for them, or who would have given up because their submissions kept failing for mysterious reasons?\n\nIncluding the trained checkpoint was a judgment call. Yes, it means a bunch of people will now have a meaningless score that is cluttering up the leaderboard (for a short while). But it also proves that the kernel works. If I had included an untrained checkpoint that scores 0.6931 or worse, is that because of a bug in the kernel or because the model is just bad?\n\nAnyway, not writing this to argue endlessly, just to give an insight into why I decided to publish this kernel.",
    "729755": "All fair points! And congratulations on your kernel, too!",
    "729887": "Well, to be honest, I couldn't care less if people copy pasted some kernel and got their name on the top 50 LB. But this comment is to primarily reply to your question - \"Should it only be accessible to people with good coding skills?\"\n\nIn my opinion, not having a solution forces people to improve their coding skills and therefore take on challenges which they otherwise wouldn't. I personally am not the best coder around. However, because of this competition, I have pretty much worked for long hours making my own pipeline from scratch and it feels great. Sure, I haven't even submitted a kernel yet, or even made an attempt to. But even without doing so, I have made so much personal progress on my skill level.",
    "729898": "...posted earlier in wrong thread\n    My main objection to the Kernel was that it was dependent on a training checkpoint that was provided by the author without the code that produced it. Surely a more poorly performing model could have been supplied to preserve the LB. It seemed to me that this guaranteed a large cohort that would run it and not improve on it, much like a leak (as has been mentioned). I thought that was a bad idea. But now that HA (with significant effort) has provided the training kernel, that objection is no longer there so I'm now fine with the kernel. I have myself benefited in the past from good kernels that I went on to improve.\n    I won't run it (because I like my own code, poorly performing as it is) but I will read it carefully and use the good stuff.",
    "730019": "Oversharing is a recurrent problem on Kaggle...",
    "730489": "I want to frame my words here as respectfully as I can, I am pretty sure @humananalog did this in good spirit, however also want to voice my disagreement to the sharing the original model.\n\nin one hand I am very much against sharing high scoring kernel - let's, for now, define \"high-scoring\" as top 5% LB - when a competition is well underway. this one has been ongoing for more than 6 weeks, for me losing of the  LB at this point as a reflective mechanism take too much of the fun out of the competition - it is not good for kaggle as a platform if we encounter this almost every competition. we also should take into consideration that this competition already survived a leak\n\nadditionally, I especially dislike sharing of high scoring pre-train model when there is no way to reproduce. I understand that @humananalog has already shared a training kernel with a simplified version, I still take the view that few would be able to replicate his shared model. I agree with @feifeizaici that this is too high a score for a baseline model at the current point. Time will tell, I am prepared to be proven wrong on this one. \n\nIn the other hand, I really appreciate @humananalog effort in making the inference pipeline available, and the subsequent effort publishes part of his training scheme - nice touches, and I personally am learning from his work. \n\nso that is it, I would stop bitching about it in this competition forum, and suffer more to improve my own modelling pipeline.",
    "730509": "Thank for sharing bro, I'm just a beginner and I learn a lot from this site. I'm just coming up with simple ideas to implement, sometime people come up with complex kernels and you don't learn as much if you just go through and implement a simple model.\n\nWill go through your submissions today. Thanks for this thread, I wouldn't have noticed it otherwise",
    "731313": "Agree, he took the fun in the game.",
    "731428": "Yet, only three days after I posted the inference kernel, the people who used it are quickly dropping off the first page of the leaderboard... So it doesn't seem to have done too much harm (and perhaps has even encouraged people to improve their solutions, who knows).",
    "731473": "Not sure I agree with your conclusion.\n\nMy conclusion would be that folks found ways to improve your kernel.   Looking at the top 100 my bet would be those folks with less than 5 submissions are mostly using your kernel.",
    "731479": "To put things in perspective a little, 96 people currently have your kernel's score. PC Jimmy here, without them, would be in 59th place rather than 155th, if we assume their scores would've otherwise been lower than his.\n\nThis corresponds with him going from silver medal to no medal.",
    "731489": "humananalog in the spirit of sharing. Here are some frames. https://www.kaggle.com/sciarrilli/dfdc-f150 \nthanks for the notebooks by the way. great for learning.",
    "731533": "Valid point, if it were the last day of the competition. Three days ago, running my kernel got you to position 30. Now it's position 75. By the end of the week, I expect those people to be out of the silver medals too.",
    "731535": "humananalog That's the case for me. I had a look to your model and tried it because my model (with similar approach) did not work so well for me (LB 0.51). Then, it showed me that my model was working fine indeed with some little adjustments (LB 0.45). One difficulty is to understand why such model overfits so quickly. We're far from the end of the competition, IMO your kernel gave a boost to force everyone to improve.",
    "731620": "Sharing leads to more/faster progress, and it likely means a better overall logloss when the final models are submitted. I don't see it as a problem.",
    "731630": "James Howard - actually if Human Analog had not shared his kernel I probably would be in IDGAF land working on a different competition.   After the first month of just getting a submission to finally work I was ready to call it quits as I realized the difficulties I was facing getting ffmpeg to work the way I wanted.  His kernel showed me a different set of tools that mesh nicely with the things I wanted to do using ffmpeg/mtcnn.  I am back to having hope that I can learn and generate new code every day.\n\nAs I type this we are sitting on 64 days left - all the folks who are throwing shade about his share - I don't recall a single competition I have worked on over the last year were we would have been sitting at starter kits for as long as this one did.  For  most of that past it seemed like at around the 30 days left to go;  was the point when shares started drying up.  This competition seemed to have dried up in the first week until his share.  \n\nSo with his share I am sitting at 155 and without his share I would be sitting on the couch watching the Impeachment of President Trump.  He saved from flipping between CNN AND Fox News trying to compute the average.",
    "731726": "Sharing is not a problem. Oversharing is.",
    "731758": "I'm very thankful to Human Analog for sharing this kernel. I've been meaning to learn Pytorch for months now, and his kernel gave me the opportunity to re-open that pytorch book, and learn everything I've been delaying. I'm not here to win competitions, I'm here to learn. Thanks for sharing !",
    "733111": "Human Analog had a goodwill to share.\nBut,  the process of deep learning is veiled and somewhat ambiguous yet.\n\nI experienced that some models work well, on the other hand other models didn't.\nAnd I can't explain it clearly.\nSo, I think kaggle competitions include all that painful labor.\n\nThere exists grey-zone between encourage to compete and learn.\nBut I think sharing the pre-trained weights scoring 100th in LB was too far \nalthough it left more than a month to the end of DFDC.\n\nIt's not pleasant to see just-copy-and-shiftenter player beat participants who\nhave been made big effort over than a month.\n\nI had been worse than 0.46788 for a long time and many participants on the way to\nimprove their model but now just stall on that score.\n\nRank is Rank.\nIt's more than a copy and paste.",
    "733326": "humananalog Thank you for your generous contributions to this competition. It's interesting to learn that the model in the inference demo was only trained using ~30000 images. Any chance you can tell us if some of these images were sampled from compressed videos?",
    "733643": "I did say ~30000 images but I just looked at the code for that particular model and it's not entirely correct. This was for an older version of the model that performed in the 0.55xxx range. \n\nThe training set I used for the 0.46xxx model has 16815 real videos and 85873 fake videos. In each batch, I sample half the batch from the ~17K real videos. The other half is sampled from the ~86K fake videos. So this actually uses ~100K images.\n\nPart of the image augmentation I do involves randomly compressing the image as JPEG with a randomly chosen quality setting. This approximates video compression.",
    "733677": "&gt; **Human Analog wrote:**\n&gt; The training set I used for the 0.46xxx model has 16815 real videos and 85873 fake videos. In each batch, I sample half the batch from the ~17K real videos. The other half is sampled from the ~86K fake videos. So this actually uses ~100K images.\n&gt; \n&gt; Part of the image augmentation I do involves randomly compressing the image as JPEG with a randomly chosen quality setting. This approximates video compression.\n\nThanks for making this clarification, on my side with my own dataset, my observation is that I can get to about 0.51LB with 50k images, and probably to a similar level to 0.45LB with 100k - I haven't check for models using this specific amount of training data. \n\nIt is reassuring to see other's LB score and training data amount is somehow correlate to our own.",
    "733692": "My biggest gain came from combining the real and fake images inside the same batch, so that the model always sees a balanced amount. (This doesn't choose the fake images randomly, but samples from the fakes that belong to the reals that have been chosen from the batch.)",
    "733695": "interesting, and thanks for sharing. \nfor me the largest gain so far come from 1) increasing amount of training data to a certain level, and 2) getting augmentation right - i.e. not too heavy and not too light\n\nand of course, having a reliable inference pipleline is also critical - I am sure we agree on this one 😉",
    "734286": "",
    "734518": "sometimes I train the whole network, sometime I do some fine-tune before training the whole network - my ad-hoc observation seems to suggest that this works better with smaller training size, \naside from that, it really depends on my mood.... :)",
    "734776": "humananalog Awesome, thanks for the advice",
    "748382": "That's what happens when sharing is gamified :)"
  },
  "source": "meta"
}