{
  "id": 124728,
  "title": "CV vs LB",
  "url": "/competitions/deepfake-detection-challenge/discussion/124728",
  "author_name": "YoonSoo",
  "post_date": "2020-01-06T07:34:57.032000",
  "votes": 34,
  "comment_count": 80,
  "views": 0,
  "content": "<p>These are validation strategies I've employed.</p>\n\n<ol>\n<li>Split train set and validation set with respect to source videos (see <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/124127\">here</a>)</li>\n<li>Make positive-negative ratio to 0.5 : 0.5 in validation set</li>\n<li>Simulated 1/4 resolution augmentations to 2/9 of the validation videos (see <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122013\">here</a>)</li>\n</ol>\n\n<p>And these are the scores I got.</p>\n\n<ul>\n<li>validation logloss: around 0.20</li>\n<li>public leaderboard logloss: around 0.60</li>\n</ul>\n\n<p>I'm suspecting three causes why the scores have such a large discrepancy.</p>\n\n<ol>\n<li>My inference notebook has some bug in it. (I've checked it thoroughly, but it is possible.)</li>\n<li>As discussed <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/124417\">here</a>, there may exist duplicates in train data, and I probably got duplicate videos both in train set and validation set.</li>\n<li>Maybe test set is comprised only of actors that don't appear in provided train set, and my model overfitted extremely to certain face identities.</li>\n</ol>\n\n<p>What is your validation score and public leaderboard score? (without leakage) Also, I'd appreciate it if you guys share your thoughts about this gap between scores.</p>\n\n<p><em>Update) I fixed some bug in my code (I was new to pytorch and I had some issues), added some augmentations, etc. Just using folder split, not using augmentations for validation set. -&gt; Now validation logloss 0.22770, public logloss 0.39295</em></p>",
  "messages": [
    {
      "id": 711501,
      "postDate": "2020-01-06T07:34:57.033Z",
      "content": "<p>These are validation strategies I've employed.</p>\n\n<ol>\n<li>Split train set and validation set with respect to source videos (see <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/124127\">here</a>)</li>\n<li>Make positive-negative ratio to 0.5 : 0.5 in validation set</li>\n<li>Simulated 1/4 resolution augmentations to 2/9 of the validation videos (see <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122013\">here</a>)</li>\n</ol>\n\n<p>And these are the scores I got.</p>\n\n<ul>\n<li>validation logloss: around 0.20</li>\n<li>public leaderboard logloss: around 0.60</li>\n</ul>\n\n<p>I'm suspecting three causes why the scores have such a large discrepancy.</p>\n\n<ol>\n<li>My inference notebook has some bug in it. (I've checked it thoroughly, but it is possible.)</li>\n<li>As discussed <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/124417\">here</a>, there may exist duplicates in train data, and I probably got duplicate videos both in train set and validation set.</li>\n<li>Maybe test set is comprised only of actors that don't appear in provided train set, and my model overfitted extremely to certain face identities.</li>\n</ol>\n\n<p>What is your validation score and public leaderboard score? (without leakage) Also, I'd appreciate it if you guys share your thoughts about this gap between scores.</p>\n\n<p><em>Update) I fixed some bug in my code (I was new to pytorch and I had some issues), added some augmentations, etc. Just using folder split, not using augmentations for validation set. -&gt; Now validation logloss 0.22770, public logloss 0.39295</em></p>",
      "rawMarkdown": "These are validation strategies I've employed.\n\n1. Split train set and validation set with respect to source videos (see [here](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/124127))\n2. Make positive-negative ratio to 0.5 : 0.5 in validation set\n3. Simulated 1/4 resolution augmentations to 2/9 of the validation videos (see [here](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122013))\n\nAnd these are the scores I got.\n\n* validation logloss: around 0.20\n* public leaderboard logloss: around 0.60\n\nI'm suspecting three causes why the scores have such a large discrepancy.\n\n1. My inference notebook has some bug in it. (I've checked it thoroughly, but it is possible.)\n2. As discussed [here](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/124417), there may exist duplicates in train data, and I probably got duplicate videos both in train set and validation set.\n3. Maybe test set is comprised only of actors that don't appear in provided train set, and my model overfitted extremely to certain face identities.\n\nWhat is your validation score and public leaderboard score? (without leakage) Also, I'd appreciate it if you guys share your thoughts about this gap between scores.\n\n*Update) I fixed some bug in my code (I was new to pytorch and I had some issues), added some augmentations, etc. Just using folder split, not using augmentations for validation set. -&gt; Now validation logloss 0.22770, public logloss 0.39295*",
      "votes": 34
    },
    {
      "id": 738174,
      "postDate": "2020-02-06T08:32:17.423Z",
      "content": "<p>Limiting it to my best kernel using a single model:\nValidation loss: 0.1887\nLeaderboard loss: 0.39309</p>",
      "rawMarkdown": "Limiting it to my best kernel using a single model:\nValidation loss: 0.1887\nLeaderboard loss: 0.39309",
      "votes": 9,
      "replies": [
        {
          "id": 738180,
          "postDate": "2020-02-06T08:37:18.637Z",
          "content": "<p>Hi James, the \"Validation loss: 0.1887\" is the score for one fold or the average score of k-fold?</p>",
          "rawMarkdown": "Hi James, the \"Validation loss: 0.1887\" is the score for one fold or the average score of k-fold?",
          "votes": 2
        },
        {
          "id": 738186,
          "postDate": "2020-02-06T08:40:25.720Z",
          "content": "<p>I only train a single fold. I do a single train/val split, chunks 30-39 being the validation set (arbitrarily).</p>",
          "rawMarkdown": "I only train a single fold. I do a single train/val split, chunks 30-39 being the validation set (arbitrarily).",
          "votes": 10
        },
        {
          "id": 738198,
          "postDate": "2020-02-06T08:58:10.077Z",
          "content": "<p>Thanks for your information. It's really useful for my team!</p>",
          "rawMarkdown": "Thanks for your information. It's really useful for my team!"
        },
        {
          "id": 738223,
          "postDate": "2020-02-06T09:29:03.647Z",
          "content": "<p><a href=\"/jamesphoward\">@jamesphoward</a> May I ask why you chose chunks 30-39 as the validation set?</p>",
          "rawMarkdown": "@jamesphoward May I ask why you chose chunks 30-39 as the validation set?"
        },
        {
          "id": 738230,
          "postDate": "2020-02-06T09:39:09.860Z",
          "content": "<p>Completely randomly, as all things should be :)</p>\n\n<p>But really, rationale was I wanted a continuous chunk, hoping that'd reduce actor/algorithm overlap as much as possible. And I picked something middling in case there were a couple of strategies they'd used to create the videos and ones in the later chunks may have been different algorithms from the start.</p>",
          "rawMarkdown": "Completely randomly, as all things should be :)\n\nBut really, rationale was I wanted a continuous chunk, hoping that'd reduce actor/algorithm overlap as much as possible. And I picked something middling in case there were a couple of strategies they'd used to create the videos and ones in the later chunks may have been different algorithms from the start.",
          "votes": 1
        },
        {
          "id": 738234,
          "postDate": "2020-02-06T09:52:17.723Z",
          "content": "<blockquote>\n  <p><strong>James Howard wrote:</strong></p>\n  \n  <p>Limiting it to my best kernel using a single model:\n  Validation loss: 0.1887\n  Leaderboard loss: 0.39309</p>\n</blockquote>\n\n<p>looks like some wonderful engineering work has been put in place to stack things up to give you the additional boost on LB 💯 </p>",
          "rawMarkdown": "&gt; **James Howard wrote:**\n&gt; \n&gt; Limiting it to my best kernel using a single model:\n&gt; Validation loss: 0.1887\n&gt; Leaderboard loss: 0.39309\n\nlooks like some wonderful engineering work has been put in place to stack things up to give you the additional boost on LB 💯 ",
          "votes": 1
        },
        {
          "id": 738238,
          "postDate": "2020-02-06T09:56:38.003Z",
          "content": "<p><a href=\"/jamesphoward\">@jamesphoward</a>  If you don't mind me asking, what resolution of images are you working with?</p>",
          "rawMarkdown": "@jamesphoward  If you don't mind me asking, what resolution of images are you working with?"
        },
        {
          "id": 738240,
          "postDate": "2020-02-06T10:01:57.917Z",
          "content": "<p>I resize bits of interest to 256 x 256, with 0padding to make square, and then center-cropp to 224 x 224.</p>",
          "rawMarkdown": "I resize bits of interest to 256 x 256, with 0padding to make square, and then center-cropp to 224 x 224.",
          "votes": 6
        },
        {
          "id": 772161,
          "postDate": "2020-03-15T05:42:36.380Z",
          "content": "<p><a href=\"/jamesphoward\">@jamesphoward</a>  What are do approximately? input for CNN is single face? Custom or public architecture?</p>",
          "rawMarkdown": "@jamesphoward  What are do approximately? input for CNN is single face? Custom or public architecture?"
        }
      ]
    },
    {
      "id": 713013,
      "postDate": "2020-01-07T20:14:40.837Z",
      "content": "<p><a href=\"/harangdev\">@harangdev</a> , I did something similar, with the following tweaks:\n- 1/3 of the videos I kept unchanged\n- 2/9 of the videos I resized to 1/4 of their sizes\n- 2/9 of the videos I reduced FPS to 15\n- 2/9 of the videos I applied a hard compression</p>\n\n<p>This is exactly what is described in DFDC paper.</p>\n\n<p>My public LB is about the same as yours (~0.63), but my validation logloss is ~0.55, much closer to my LB.\nI suspect the key is the last bullet: apply a hard compression. This reduces the videos' file sizes to &lt;1/10 of their original sizes, and make it much harder for our algos to correctly classify as fake or real.\nIMPORTANT: I made sure these 4 proportions are respected in both training and validation sets.</p>",
      "rawMarkdown": "@harangdev , I did something similar, with the following tweaks:\n- 1/3 of the videos I kept unchanged\n- 2/9 of the videos I resized to 1/4 of their sizes\n- 2/9 of the videos I reduced FPS to 15\n- 2/9 of the videos I applied a hard compression\n\nThis is exactly what is described in DFDC paper.\n\nMy public LB is about the same as yours (~0.63), but my validation logloss is ~0.55, much closer to my LB.\nI suspect the key is the last bullet: apply a hard compression. This reduces the videos' file sizes to &lt;1/10 of their original sizes, and make it much harder for our algos to correctly classify as fake or real.\nIMPORTANT: I made sure these 4 proportions are respected in both training and validation sets.",
      "votes": 7,
      "replies": [
        {
          "id": 713244,
          "postDate": "2020-01-08T04:20:31.273Z",
          "content": "<p>Yeah, I will implement compression augmentation later, but I'm not sure it will have impact on my validation score that much, like from 0.2 to 0.6. Thanks for your suggestions!</p>",
          "rawMarkdown": "Yeah, I will implement compression augmentation later, but I'm not sure it will have impact on my validation score that much, like from 0.2 to 0.6. Thanks for your suggestions!"
        },
        {
          "id": 714298,
          "postDate": "2020-01-09T09:26:09.903Z",
          "rawMarkdown": ""
        },
        {
          "id": 733394,
          "postDate": "2020-01-31T04:14:21.730Z",
          "content": "<p>What exactly do you mean by hard compression? in albumentations we have an option for jpeg compression. Is that similar to what you used?</p>",
          "rawMarkdown": "What exactly do you mean by hard compression? in albumentations we have an option for jpeg compression. Is that similar to what you used?"
        },
        {
          "id": 733407,
          "postDate": "2020-01-31T04:50:17.293Z",
          "content": "<p>I'm also searching for the same. until now the best I can think to reduce the overall quality is to resize image size to (70,70) and then resize it to (224,224). Of course, numbers are just for illustration.</p>",
          "rawMarkdown": "I'm also searching for the same. until now the best I can think to reduce the overall quality is to resize image size to (70,70) and then resize it to (224,224). Of course, numbers are just for illustration."
        },
        {
          "id": 733496,
          "postDate": "2020-01-31T08:05:40.923Z",
          "content": "<p>You can use both JpegCompression and Downscale from Albumentations:</p>\n\n<p><code>\nA.OneOf([\n    A.JpegCompression(quality_lower=8, quality_upper=30, p=1.0),\n    A.Downscale(scale_min=0.25, scale_max=0.75, p=1.0)\n ], p=0.22)\n</code></p>",
          "rawMarkdown": "You can use both JpegCompression and Downscale from Albumentations:\n\n```\nA.OneOf([\n    A.JpegCompression(quality_lower=8, quality_upper=30, p=1.0),\n    A.Downscale(scale_min=0.25, scale_max=0.75, p=1.0)\n ], p=0.22)\n```",
          "votes": 6
        },
        {
          "id": 742914,
          "postDate": "2020-02-11T16:34:05.430Z",
          "content": "<p><a href=\"/mpware\">@mpware</a> Why did you choose these Jpeg quality values? Are they related to the H.264 CRF values (e.g. default 23 corresponds to quality_upper=30)?</p>",
          "rawMarkdown": "@mpware Why did you choose these Jpeg quality values? Are they related to the H.264 CRF values (e.g. default 23 corresponds to quality_upper=30)?"
        },
        {
          "id": 743040,
          "postDate": "2020-02-11T18:27:08.587Z",
          "content": "<p>No, I've tried different settings to make JPEG compression hard.</p>",
          "rawMarkdown": "No, I've tried different settings to make JPEG compression hard.",
          "votes": 1
        }
      ]
    },
    {
      "id": 772401,
      "postDate": "2020-03-15T12:53:26.560Z",
      "content": "<p>I believe just splitting the folders 0-40, 41-49 gives a good validation method here's my results:\n```\nValidation Log Loss: .3347\nLB: .470</p>\n\n<p>Validation Log Loss: .3035\nLB: .446</p>\n\n<p>Validation Log Loss: .32555\nLB: .439\n<code>\nUPDATE:\nWith 0-39, 40-49 folder validation:\n</code>\nVal: .259\nLB: .392\n```\nThe trick is to save a checkpoint when the model is just about to overfit which is difficult. At best you can save every checkpoint and then only submit models that give a even 50/50 split on the test 400 vids.</p>",
      "rawMarkdown": "I believe just splitting the folders 0-40, 41-49 gives a good validation method here's my results:\n```\nValidation Log Loss: .3347\nLB: .470\n\nValidation Log Loss: .3035\nLB: .446\n\nValidation Log Loss: .32555\nLB: .439\n```\nUPDATE:\nWith 0-39, 40-49 folder validation:\n```\nVal: .259\nLB: .392\n```\nThe trick is to save a checkpoint when the model is just about to overfit which is difficult. At best you can save every checkpoint and then only submit models that give a even 50/50 split on the test 400 vids.",
      "votes": 3,
      "replies": [
        {
          "id": 772421,
          "postDate": "2020-03-15T13:20:56.687Z",
          "content": "<p>My results were 0.160 on the same data, 0.318 on leaderboard.</p>",
          "rawMarkdown": "My results were 0.160 on the same data, 0.318 on leaderboard.",
          "votes": 1
        },
        {
          "id": 774974,
          "postDate": "2020-03-16T06:05:38.137Z",
          "content": "<p>Hi. I'm late joiner. <a href=\"/greatgamedota\">@greatgamedota</a> and <a href=\"/harshitsheoran\">@harshitsheoran</a> comments are really helpful. Thx!</p>",
          "rawMarkdown": "Hi. I'm late joiner. @greatgamedota and @harshitsheoran comments are really helpful. Thx!"
        },
        {
          "id": 775166,
          "postDate": "2020-03-16T11:24:02.767Z",
          "content": "<p><a href=\"/harshitsheoran\">@harshitsheoran</a> Could you please share how you construct the validation set? I've tried a lot of strategies, but there's always a big gap between CV and LB.</p>",
          "rawMarkdown": "@harshitsheoran Could you please share how you construct the validation set? I've tried a lot of strategies, but there's always a big gap between CV and LB."
        },
        {
          "id": 775213,
          "postDate": "2020-03-16T12:18:46.373Z",
          "content": "<p>Well, <a href=\"/chenshen03\">@chenshen03</a> it looks to me, that my validation gap is the largest, its double score  on leaderboard every time.</p>\n\n<p>Anyways, first list 10 last folders videos in a format of filenames, originals, then randomly take 500 of those, which are from my experience, testing was same for me as testing all folders all videos or just 500, which are chosen by shuffle([filenames,original], random_state=42)[:500] where shuffle is sklearn.utils.shuffle.</p>",
          "rawMarkdown": "Well, @chenshen03 it looks to me, that my validation gap is the largest, its double score  on leaderboard every time.\n\nAnyways, first list 10 last folders videos in a format of filenames, originals, then randomly take 500 of those, which are from my experience, testing was same for me as testing all folders all videos or just 500, which are chosen by shuffle([filenames,original], random_state=42)[:500] where shuffle is sklearn.utils.shuffle."
        }
      ]
    },
    {
      "id": 712174,
      "postDate": "2020-01-06T22:47:08.167Z",
      "content": "<p>My cv was 0.667 and my LB was 0.69243.</p>",
      "rawMarkdown": "My cv was 0.667 and my LB was 0.69243.",
      "votes": 3,
      "replies": [
        {
          "id": 712182,
          "postDate": "2020-01-06T23:03:54.747Z",
          "content": "<p>Thank you for sharing. Your scores align pretty well. How did you split validation set?</p>",
          "rawMarkdown": "Thank you for sharing. Your scores align pretty well. How did you split validation set?"
        },
        {
          "id": 712202,
          "postDate": "2020-01-07T00:00:45.207Z",
          "content": "<p>I just found out a glitch in my validation strategy. I passed the string labels instead of number labels. The new CV is 0.69312.\nI split the validation set by folders(last two folders are the validation set).</p>",
          "rawMarkdown": "I just found out a glitch in my validation strategy. I passed the string labels instead of number labels. The new CV is 0.69312.\nI split the validation set by folders(last two folders are the validation set).",
          "votes": 2
        },
        {
          "id": 712213,
          "postDate": "2020-01-07T00:23:14.360Z",
          "content": "<p>Thanks for your reply. I think I should investigate further into my data splits.</p>",
          "rawMarkdown": "Thanks for your reply. I think I should investigate further into my data splits.",
          "votes": 1
        },
        {
          "id": 712217,
          "postDate": "2020-01-07T00:30:18.297Z",
          "content": "<p>Can you post the code here as well(if you don't feel like making the whole notebook public, just publish the key parts or replace the parts you want to keep secret with a brief summary of that part)? Thanks!</p>",
          "rawMarkdown": "Can you post the code here as well(if you don't feel like making the whole notebook public, just publish the key parts or replace the parts you want to keep secret with a brief summary of that part)? Thanks!"
        },
        {
          "id": 712239,
          "postDate": "2020-01-07T01:26:01.580Z",
          "content": "<p>```</p>\n\n<h1>load metadata to variable meta</h1>\n\n<p>meta['label'] = meta['label'].map({'FAKE': 1, 'REAL': 0})\nmeta = meta.set_index('filename')</p>\n\n<p>def split(meta, pos_val_size=2000):</p>\n\n<pre><code>meta = meta.sample(frac=1, random_state=0)\nsources = meta['original'].dropna().unique()\nvalid_pos = sources[:pos_val_size]\ntmp = meta.drop(valid_pos).reset_index().set_index('original')\nvalid_neg = tmp.loc[valid_pos, 'index'].groupby(level=0, group_keys=False).apply(lambda x: x.sample(1, random_state=0)).values\nvalid = np.concatenate([valid_pos, valid_neg])\ntrain_pos = sources[pos_val_size:]\ntrain_neg = tmp.loc[sources[train_pos], 'index'].values\ntrain = np.concatenate([train_pos, train_neg])\n\nnp.random.seed(0)\nnp.random.shuffle(train)\nnp.random.seed(0)\nnp.random.shuffle(valid)\ntrain = [x[:-4] for x in train]\nvalid = [x[:-4] for x in valid]\n\nreturn train, valid\n</code></pre>\n\n<p>train, valid = split(meta)\n```</p>",
          "rawMarkdown": "```\n# load metadata to variable meta\nmeta['label'] = meta['label'].map({'FAKE': 1, 'REAL': 0})\nmeta = meta.set_index('filename')\n\ndef split(meta, pos_val_size=2000):\n    \n    meta = meta.sample(frac=1, random_state=0)\n    sources = meta['original'].dropna().unique()\n    valid_pos = sources[:pos_val_size]\n    tmp = meta.drop(valid_pos).reset_index().set_index('original')\n    valid_neg = tmp.loc[valid_pos, 'index'].groupby(level=0, group_keys=False).apply(lambda x: x.sample(1, random_state=0)).values\n    valid = np.concatenate([valid_pos, valid_neg])\n    train_pos = sources[pos_val_size:]\n    train_neg = tmp.loc[sources[train_pos], 'index'].values\n    train = np.concatenate([train_pos, train_neg])\n\n    np.random.seed(0)\n    np.random.shuffle(train)\n    np.random.seed(0)\n    np.random.shuffle(valid)\n    train = [x[:-4] for x in train]\n    valid = [x[:-4] for x in valid]\n    \n    return train, valid\n\ntrain, valid = split(meta)\n```",
          "votes": 1
        }
      ]
    },
    {
      "id": 714028,
      "postDate": "2020-01-09T00:28:14.537Z",
      "content": "<p>I've got cv:0.49 lb:0.53</p>\n\n<p>Apparently my strategy is working well enough as before I had cv:0.55/lb:0.59, so its tracking well.\nI'm using the strategy to split by file and downsample train and validation so it has roughly 50/50 fake and real.\nStill not using all the data available and doing a similar augmentation as the one <a href=\"/carlossouza\">@carlossouza</a> suggested here.</p>",
      "rawMarkdown": "I've got cv:0.49 lb:0.53\n\nApparently my strategy is working well enough as before I had cv:0.55/lb:0.59, so its tracking well.\nI'm using the strategy to split by file and downsample train and validation so it has roughly 50/50 fake and real.\nStill not using all the data available and doing a similar augmentation as the one @carlossouza suggested here.",
      "votes": 4,
      "replies": [
        {
          "id": 714036,
          "postDate": "2020-01-09T00:58:36.507Z",
          "content": "<p>Your scores align pretty well! I'm thinking towards that my local code has some bug in it.</p>",
          "rawMarkdown": "Your scores align pretty well! I'm thinking towards that my local code has some bug in it."
        },
        {
          "id": 724330,
          "postDate": "2020-01-21T04:00:48.180Z",
          "content": "<p><a href=\"/pedromb\">@pedromb</a>  May I ask how you organize your validation set?  For example, fetch part 1,2,3,4,5 as validation or something like that?</p>",
          "rawMarkdown": "@pedromb  May I ask how you organize your validation set?  For example, fetch part 1,2,3,4,5 as validation or something like that?"
        },
        {
          "id": 724622,
          "postDate": "2020-01-21T10:03:39.230Z",
          "content": "<p>I create a dataframe with the metadata and include a column that specify the part, then I use sklearn <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GroupShuffleSplit.html\">GroupShuffleSplit</a>.</p>\n\n<p>I just set the train size in a way that my test size has at least 2k real videos.</p>",
          "rawMarkdown": "I create a dataframe with the metadata and include a column that specify the part, then I use sklearn [GroupShuffleSplit](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GroupShuffleSplit.html).\n\nI just set the train size in a way that my test size has at least 2k real videos.",
          "votes": 3
        },
        {
          "id": 735365,
          "postDate": "2020-02-02T23:36:10.333Z",
          "content": "<p><a href=\"/pedromb\">@pedromb</a> do you think that the test dataset is 50/50 fake/real, so you split the same on validation set? \nin Train set, i use weighted sampling 50/50 get CV: 0.4\nI keep the original ratio 10/50 real/fake in validation set, get much better CV result: 0.2\nDoes it infer that test set distrubtion is closet to 50/50\nThanks</p>",
          "rawMarkdown": "@pedromb do you think that the test dataset is 50/50 fake/real, so you split the same on validation set? \nin Train set, i use weighted sampling 50/50 get CV: 0.4\nI keep the original ratio 10/50 real/fake in validation set, get much better CV result: 0.2\nDoes it infer that test set distrubtion is closet to 50/50\nThanks"
        },
        {
          "id": 735395,
          "postDate": "2020-02-03T01:12:29.223Z",
          "content": "<p>The test set is definitely 50:50.</p>\n\n<p><a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/126524\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/126524</a></p>",
          "rawMarkdown": "The test set is definitely 50:50.\n\nhttps://www.kaggle.com/c/deepfake-detection-challenge/discussion/126524"
        }
      ]
    },
    {
      "id": 735496,
      "postDate": "2020-02-03T05:35:19.333Z",
      "content": "<p>EDIT:\ncv:0.0978 lb:0.0.40492</p>",
      "rawMarkdown": "EDIT:\ncv:0.0978 lb:0.0.40492",
      "votes": 2,
      "replies": [
        {
          "id": 735526,
          "postDate": "2020-02-03T06:43:34.913Z",
          "content": "<p>You're doing great at cv. Seems like there are more audio manipulated videos in lb dataset. Can I know how you made the test-val split?</p>",
          "rawMarkdown": "You're doing great at cv. Seems like there are more audio manipulated videos in lb dataset. Can I know how you made the test-val split?"
        },
        {
          "id": 736205,
          "postDate": "2020-02-04T00:15:27.503Z",
          "content": "<p>0.15 test_rate simple sklearn train_test_split. I will indeed try clustering for more accurate validation score.</p>",
          "rawMarkdown": "0.15 test_rate simple sklearn train_test_split. I will indeed try clustering for more accurate validation score.",
          "votes": 1
        },
        {
          "id": 736426,
          "postDate": "2020-02-04T07:01:04.580Z",
          "content": "<p>I'm using first and last chuncks for validation (cv0.43, lb0.47). I think this helps due to variations in first set and audio manipulated videos in last set. Please share the results after clustering. </p>",
          "rawMarkdown": "I'm using first and last chuncks for validation (cv0.43, lb0.47). I think this helps due to variations in first set and audio manipulated videos in last set. Please share the results after clustering. "
        },
        {
          "id": 738107,
          "postDate": "2020-02-06T06:21:30.080Z",
          "content": "<p><a href=\"/unkownhihi\">@unkownhihi</a> Is your CV loss based on image or video?</p>",
          "rawMarkdown": "@unkownhihi Is your CV loss based on image or video?"
        }
      ]
    },
    {
      "id": 733319,
      "postDate": "2020-01-31T00:35:44.837Z",
      "content": "<p><a href=\"/harangdev\">@harangdev</a>  I am wondering, what was the fix for your validation logloss? Was it the augmentations? \nI am also experiencing the discrepancy between the two.</p>",
      "rawMarkdown": "@harangdev  I am wondering, what was the fix for your validation logloss? Was it the augmentations? \nI am also experiencing the discrepancy between the two.",
      "votes": 2,
      "replies": [
        {
          "id": 733321,
          "postDate": "2020-01-31T00:54:06.947Z",
          "content": "<p>I haven't figure out how to make the cv and lb align yet.☹️</p>",
          "rawMarkdown": "I haven't figure out how to make the cv and lb align yet.☹️",
          "votes": 2
        }
      ]
    },
    {
      "id": 714324,
      "postDate": "2020-01-09T09:44:21.470Z",
      "content": "<p>cv:0.48 lb:0.50, which looks normal</p>",
      "rawMarkdown": "cv:0.48 lb:0.50, which looks normal\n\n",
      "votes": 2,
      "replies": [
        {
          "id": 714443,
          "postDate": "2020-01-09T12:52:59.600Z",
          "content": "<p>That's nice. Did you use full data?</p>",
          "rawMarkdown": "That's nice. Did you use full data?"
        },
        {
          "id": 714477,
          "postDate": "2020-01-09T13:30:18.140Z",
          "content": "<p>Almost the full data, I split apart some videos to try other methods.</p>",
          "rawMarkdown": "Almost the full data, I split apart some videos to try other methods.",
          "votes": 2
        },
        {
          "id": 715009,
          "postDate": "2020-01-10T02:57:42.163Z",
          "content": "<p><a href=\"/feifeizaici\">@feifeizaici</a> May I ask you how you split validation data? I'm still getting too good validation scores even after trying dozen experiments.</p>",
          "rawMarkdown": "@feifeizaici May I ask you how you split validation data? I'm still getting too good validation scores even after trying dozen experiments.",
          "votes": 2
        },
        {
          "id": 715012,
          "postDate": "2020-01-10T03:10:59.197Z",
          "content": "<p>I colleced 2000 FAKE and 2000 REAL videos, and for each of them I extracted 5 frames. My validation set is right these 20k faces. I think val/test images and train images should not come from a same batch of videos, for some frames are just too similar. </p>",
          "rawMarkdown": "I colleced 2000 FAKE and 2000 REAL videos, and for each of them I extracted 5 frames. My validation set is right these 20k faces. I think val/test images and train images should not come from a same batch of videos, for some frames are just too similar. ",
          "votes": 5
        },
        {
          "id": 715970,
          "postDate": "2020-01-11T04:13:58.077Z",
          "content": "<p><a href=\"/feifeizaici\">@feifeizaici</a> cv: 0.31, lb: 0.75. My model is overfitting. Do you have any data enhancement techniques? \nIn addition, the longer the training, the more severe the overfitting. To what extent does your model generally train? Thank you very much!</p>",
          "rawMarkdown": "@feifeizaici cv: 0.31, lb: 0.75. My model is overfitting. Do you have any data enhancement techniques? \nIn addition, the longer the training, the more severe the overfitting. To what extent does your model generally train? Thank you very much!",
          "votes": 1
        },
        {
          "id": 715974,
          "postDate": "2020-01-11T04:20:09.040Z",
          "content": "<p>Looks like your score is a lot better than 0.75 😄 </p>",
          "rawMarkdown": "Looks like your score is a lot better than 0.75 😄 "
        },
        {
          "id": 715978,
          "postDate": "2020-01-11T04:48:31.607Z",
          "content": "<p><a href=\"/unkownhihi\">@unkownhihi</a>  My current score is obtained by randomly sampling 10K FAKE and REAL videos. Later, I tried to divide the training set and validation set according to origin videos, and increased the margin of face detection, but overfitting occurred.</p>",
          "rawMarkdown": "@unkownhihi  My current score is obtained by randomly sampling 10K FAKE and REAL videos. Later, I tried to divide the training set and validation set according to origin videos, and increased the margin of face detection, but overfitting occurred."
        },
        {
          "id": 716051,
          "postDate": "2020-01-11T07:19:38.733Z",
          "content": "<p>Just random horizontal flip. \nI think you can set a learning rate decay to aviod overfit.</p>",
          "rawMarkdown": "Just random horizontal flip. \nI think you can set a learning rate decay to aviod overfit.",
          "votes": 1
        },
        {
          "id": 726548,
          "postDate": "2020-01-23T04:17:39.960Z",
          "content": "<p><a href=\"/chenshen03\">@chenshen03</a> I seem to get the same thing happening when I use origin videos. Massive overfitting. Did you ever figure it out?</p>",
          "rawMarkdown": "@chenshen03 I seem to get the same thing happening when I use origin videos. Massive overfitting. Did you ever figure it out?",
          "votes": 1
        }
      ]
    },
    {
      "id": 713239,
      "postDate": "2020-01-08T04:14:31.157Z",
      "content": "<p>CV logloss is 0.08 using 8% of the total 500gb data for training and 2% for validating, I know this is not the greastest data to train my model on and I think that it is also overfitting on the faces of some people (confident to solve it) I mean validation accuracy is 96%, only a few tricks, but my score is stuck on 0.69314 does not matter how much I change the sub, I guess it is a bug in inference, If someone can help me solve it, I might go for making some of my code and tricks public sometime later...,\nAlso, if someone is super super good and confident, willing to work almost full time every day (I am saying that cuz last time I asked for teammates I got nobbies) can join my team after a test ofc. I only have a 1080 ti and a slow internet... (20 Mbps)</p>\n\n<p>I reken that the score is not gonna be good on the leaderboard the reason is the test data is a lot different then the data we are training on, I guess we will be able to solve that with some custom augs, I have no augs in my data currently...</p>\n\n<p><a href=\"/harangdev\">@harangdev</a> Thank you for this discussion, so many ideas I have not tried...</p>",
      "rawMarkdown": "CV logloss is 0.08 using 8% of the total 500gb data for training and 2% for validating, I know this is not the greastest data to train my model on and I think that it is also overfitting on the faces of some people (confident to solve it) I mean validation accuracy is 96%, only a few tricks, but my score is stuck on 0.69314 does not matter how much I change the sub, I guess it is a bug in inference, If someone can help me solve it, I might go for making some of my code and tricks public sometime later...,\nAlso, if someone is super super good and confident, willing to work almost full time every day (I am saying that cuz last time I asked for teammates I got nobbies) can join my team after a test ofc. I only have a 1080 ti and a slow internet... (20 Mbps)\n\nI reken that the score is not gonna be good on the leaderboard the reason is the test data is a lot different then the data we are training on, I guess we will be able to solve that with some custom augs, I have no augs in my data currently...\n\n@harangdev Thank you for this discussion, so many ideas I have not tried...",
      "votes": 2,
      "replies": [
        {
          "id": 713247,
          "postDate": "2020-01-08T04:25:13.243Z",
          "content": "<p>Hope you fix the bug in your inference code soon. I'm curious what your lb score will be.</p>",
          "rawMarkdown": "Hope you fix the bug in your inference code soon. I'm curious what your lb score will be.",
          "votes": 1
        },
        {
          "id": 713253,
          "postDate": "2020-01-08T04:35:34.013Z",
          "content": "<p>Trust me this one is really gonna be bad I can feel it because which ever model I make I test it to some custom tests to get where the model is lacking which data it is not able to get right, And looking at the test set there are so many different people that if I dont train on full data it just does not make sense that model is gonna do any good then even random, also I dont have the pc to train it on full data, definately I can optimize it but, that will still take like 9 days for data creation and then 10-12 hours for each epoch that really pisses me off and I dont have the internet speed to just download the data I am downloading it every slowly ;-( basically that is the reason why I need a teammate...</p>",
          "rawMarkdown": "Trust me this one is really gonna be bad I can feel it because which ever model I make I test it to some custom tests to get where the model is lacking which data it is not able to get right, And looking at the test set there are so many different people that if I dont train on full data it just does not make sense that model is gonna do any good then even random, also I dont have the pc to train it on full data, definately I can optimize it but, that will still take like 9 days for data creation and then 10-12 hours for each epoch that really pisses me off and I dont have the internet speed to just download the data I am downloading it every slowly ;-( basically that is the reason why I need a teammate..."
        },
        {
          "id": 713331,
          "postDate": "2020-01-08T06:57:17.587Z",
          "content": "<p>You should go with google colab(with google drive) and transfer.sh I use combination of both. Transfer.sh is helpful for storing preprocessed videos and google drive for trained models. Then I have some wrote some functions to download a training set, uploading a folder to kaggle datatset, get face bounding box from a image as numpy array etc. I can provide if you're interested.</p>",
          "rawMarkdown": "You should go with google colab(with google drive) and transfer.sh I use combination of both. Transfer.sh is helpful for storing preprocessed videos and google drive for trained models. Then I have some wrote some functions to download a training set, uploading a folder to kaggle datatset, get face bounding box from a image as numpy array etc. I can provide if you're interested."
        },
        {
          "id": 713354,
          "postDate": "2020-01-08T07:28:05.580Z",
          "content": "<p>Why go with google colab???, it is slower than my pc in terms of processing power and I told you (copying this again lol) I can optimize it but, that will still take like 9 days for data creation and then 10-12 hours for each epoch that really pisses me off..., And god knows how much time on google colab...</p>\n\n<p>I have got a model for face detection which is much better than MTCNN and much faster too, the best I could find..., If you can answer why then sure I will be interested.</p>",
          "rawMarkdown": "Why go with google colab???, it is slower than my pc in terms of processing power and I told you (copying this again lol) I can optimize it but, that will still take like 9 days for data creation and then 10-12 hours for each epoch that really pisses me off..., And god knows how much time on google colab...\n\nI have got a model for face detection which is much better than MTCNN and much faster too, the best I could find..., If you can answer why then sure I will be interested."
        },
        {
          "id": 713495,
          "postDate": "2020-01-08T10:52:21.213Z",
          "content": "<p>Have you ever ploted your label's distribution in the kaggle kernel? Dose it looks normal?</p>",
          "rawMarkdown": "Have you ever ploted your label's distribution in the kaggle kernel? Dose it looks normal?"
        },
        {
          "id": 713497,
          "postDate": "2020-01-08T10:54:26.140Z",
          "content": "<p><a href=\"/feifeizaici\">@feifeizaici</a> I mean I have read the submission file kernel generated, the values looks as they should be to me</p>",
          "rawMarkdown": "@feifeizaici I mean I have read the submission file kernel generated, the values looks as they should be to me"
        },
        {
          "id": 713829,
          "postDate": "2020-01-08T17:34:24.800Z",
          "content": "<p>I am not looking to join as team. I'll just help anyone having hard time to train to local setup.</p>",
          "rawMarkdown": "I am not looking to join as team. I'll just help anyone having hard time to train to local setup."
        },
        {
          "id": 713959,
          "postDate": "2020-01-08T21:13:44.250Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 714225,
          "postDate": "2020-01-09T07:52:18.167Z",
          "content": "<p>I guess I solved the problem which was making the model overfit on faces, now just 10 epoch training on the sample 400 videos, I am getting 0.4 loss on the dfdc_train_part_0 as the test set..., Someone help to solve the always same score bug. </p>",
          "rawMarkdown": "I guess I solved the problem which was making the model overfit on faces, now just 10 epoch training on the sample 400 videos, I am getting 0.4 loss on the dfdc_train_part_0 as the test set..., Someone help to solve the always same score bug. ",
          "votes": 1
        },
        {
          "id": 779819,
          "postDate": "2020-03-19T17:52:14.727Z",
          "content": "<p>May I know, what was the same score bug?</p>",
          "rawMarkdown": "May I know, what was the same score bug?"
        }
      ]
    },
    {
      "id": 712269,
      "postDate": "2020-01-07T02:44:12.307Z",
      "content": "<p>cv: 0.426  LB:17.26938 </p>",
      "rawMarkdown": "cv: 0.426  LB:17.26938 ",
      "votes": 2,
      "replies": [
        {
          "id": 712277,
          "postDate": "2020-01-07T02:57:12.570Z",
          "content": "<p>LB of 17.26938 is not normal. It can be solved by infering not over the sample_sumbission videos but over the /testvideos/* files. refer to <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/123260\">link</a>\nBTW cv of 0.426 is pretty good.</p>",
          "rawMarkdown": "LB of 17.26938 is not normal. It can be solved by infering not over the sample_sumbission videos but over the /testvideos/* files. refer to [link](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/123260)\nBTW cv of 0.426 is pretty good.",
          "votes": 2
        },
        {
          "id": 712284,
          "postDate": "2020-01-07T03:03:53.870Z",
          "content": "<p>Nice to know 0.426 is good. I am trying to fixe submission file, will update.</p>",
          "rawMarkdown": "Nice to know 0.426 is good. I am trying to fixe submission file, will update."
        },
        {
          "id": 712287,
          "postDate": "2020-01-07T03:19:04.893Z",
          "content": "<p>Yeah! The best CV I can get is 0.666(unless overfit, hopefully your not overfitting). Its really not convenient that cuda can't produce reproducible results. My CV was good, but when I submit the notebook, it will train again and got bad results.  </p>",
          "rawMarkdown": "Yeah! The best CV I can get is 0.666(unless overfit, hopefully your not overfitting). Its really not convenient that cuda can't produce reproducible results. My CV was good, but when I submit the notebook, it will train again and got bad results.  ",
          "votes": 1
        },
        {
          "id": 713297,
          "postDate": "2020-01-08T06:02:15.013Z",
          "content": "<p>I can't see why inference over test_videos/ helps. Could you please elaborate.</p>",
          "rawMarkdown": "I can't see why inference over test_videos/ helps. Could you please elaborate."
        },
        {
          "id": 713932,
          "postDate": "2020-01-08T20:19:09.567Z",
          "content": "<p>The sample submission won't change but the filenames used to calculate LB will change.</p>",
          "rawMarkdown": "The sample submission won't change but the filenames used to calculate LB will change."
        }
      ]
    },
    {
      "id": 711597,
      "postDate": "2020-01-06T09:41:20.423Z",
      "content": "<p>I think it is not possible to get to validation log-loss 0.2 without a bug, especially when balanced. Do you use full train data? What is your train loss?</p>",
      "rawMarkdown": "I think it is not possible to get to validation log-loss 0.2 without a bug, especially when balanced. Do you use full train data? What is your train loss?",
      "votes": 2,
      "replies": [
        {
          "id": 711608,
          "postDate": "2020-01-06T10:05:46.547Z",
          "content": "<p>I use full train data and my train loss is about 0.077.</p>",
          "rawMarkdown": "I use full train data and my train loss is about 0.077.",
          "votes": 1
        },
        {
          "id": 713191,
          "postDate": "2020-01-08T02:25:40.977Z",
          "content": "<p>Full train data is extremely imbalanced, approx 5:1 fake to real. If you don’t treat it, your model might show great results in training, but won’t generalize well as it will have a strong bias to predict as fake, thus generating a lot of false positives. Basically there are 3 ways of dealing with class imbalance: undersampling, oversampling, or weighting the loss function. If you are using Pytorch, the following link might be useful: <a href=\"https://discuss.pytorch.org/t/dealing-with-imbalanced-datasets-in-pytorch/22596\">https://discuss.pytorch.org/t/dealing-with-imbalanced-datasets-in-pytorch/22596</a></p>",
          "rawMarkdown": "Full train data is extremely imbalanced, approx 5:1 fake to real. If you don’t treat it, your model might show great results in training, but won’t generalize well as it will have a strong bias to predict as fake, thus generating a lot of false positives. Basically there are 3 ways of dealing with class imbalance: undersampling, oversampling, or weighting the loss function. If you are using Pytorch, the following link might be useful: https://discuss.pytorch.org/t/dealing-with-imbalanced-datasets-in-pytorch/22596",
          "votes": 1
        },
        {
          "id": 713195,
          "postDate": "2020-01-08T02:27:52.653Z",
          "content": "<p><a href=\"/harangdev\">@harangdev</a>  already said:</p>\n\n<blockquote>\n  <p>Make positive-negative ratio to 0.5 : 0.5 in validation set\n  So it is balanced.</p>\n</blockquote>",
          "rawMarkdown": "@harangdev  already said:\n&gt;Make positive-negative ratio to 0.5 : 0.5 in validation set\nSo it is balanced.",
          "votes": 3
        }
      ]
    },
    {
      "id": 711540,
      "postDate": "2020-01-06T07:59:34.313Z",
      "content": "<p>Seems interesting challenge... I'll also try this problem in few days. \nThanks for sharing your Views with us.</p>",
      "rawMarkdown": "Seems interesting challenge... I'll also try this problem in few days. \nThanks for sharing your Views with us.",
      "votes": 2
    },
    {
      "id": 748515,
      "postDate": "2020-02-17T16:07:28.673Z",
      "content": "<p><a href=\"/harangdev\">@harangdev</a> Do you mind sharing your folder split number as Val set? Thanks</p>",
      "rawMarkdown": "@harangdev Do you mind sharing your folder split number as Val set? Thanks"
    },
    {
      "id": 747364,
      "postDate": "2020-02-16T10:41:54.490Z",
      "content": "<p>Some great insights here. Could you share the source of your bug, might be useful? Thanks. :) </p>",
      "rawMarkdown": "Some great insights here. Could you share the source of your bug, might be useful? Thanks. :) ",
      "replies": [
        {
          "id": 747374,
          "postDate": "2020-02-16T11:02:07.550Z",
          "content": "<p>I wasn't familiar with model.train() and model.eval() thing. Also I moved from catalyst to vanilla pytorch.</p>",
          "rawMarkdown": "I wasn't familiar with model.train() and model.eval() thing. Also I moved from catalyst to vanilla pytorch.",
          "votes": 1
        },
        {
          "id": 747408,
          "postDate": "2020-02-16T11:49:10.023Z",
          "content": "<p>Thanks!</p>",
          "rawMarkdown": "Thanks!"
        }
      ]
    },
    {
      "id": 735488,
      "postDate": "2020-02-03T05:21:04.673Z",
      "content": "<p><a href=\"/harangdev\">@harangdev</a> \nHey, if you don't mind sharing, what changes did you make to your validation splits/pipeline in order to reduce LB from 0.6 to 0.39? Looks like I'm having a very similar issue. Was it a face overfitting that was happening for you? </p>",
      "rawMarkdown": "@harangdev \nHey, if you don't mind sharing, what changes did you make to your validation splits/pipeline in order to reduce LB from 0.6 to 0.39? Looks like I'm having a very similar issue. Was it a face overfitting that was happening for you? ",
      "replies": [
        {
          "id": 738191,
          "postDate": "2020-02-06T08:49:12.617Z",
          "content": "<p>I've updated my post :)</p>",
          "rawMarkdown": "I've updated my post :)"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 738174,
      "author_name": "James Howard",
      "author_url": "",
      "post_date": "2020-02-06T08:32:17.423000",
      "content": "<p>Limiting it to my best kernel using a single model:\nValidation loss: 0.1887\nLeaderboard loss: 0.39309</p>",
      "votes": 9,
      "replies": [
        {
          "id": 738180,
          "author_name": "Duc Nguyen",
          "author_url": "",
          "post_date": "2020-02-06T08:37:18.637000",
          "content": "<p>Hi James, the \"Validation loss: 0.1887\" is the score for one fold or the average score of k-fold?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 738186,
          "author_name": "James Howard",
          "author_url": "",
          "post_date": "2020-02-06T08:40:25.720000",
          "content": "<p>I only train a single fold. I do a single train/val split, chunks 30-39 being the validation set (arbitrarily).</p>",
          "votes": 10,
          "replies": []
        },
        {
          "id": 738198,
          "author_name": "Duc Nguyen",
          "author_url": "",
          "post_date": "2020-02-06T08:58:10.077000",
          "content": "<p>Thanks for your information. It's really useful for my team!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 738223,
          "author_name": "Chason",
          "author_url": "",
          "post_date": "2020-02-06T09:29:03.647000",
          "content": "<p><a href=\"/jamesphoward\">@jamesphoward</a> May I ask why you chose chunks 30-39 as the validation set?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 738230,
          "author_name": "James Howard",
          "author_url": "",
          "post_date": "2020-02-06T09:39:09.860000",
          "content": "<p>Completely randomly, as all things should be :)</p>\n\n<p>But really, rationale was I wanted a continuous chunk, hoping that'd reduce actor/algorithm overlap as much as possible. And I picked something middling in case there were a couple of strategies they'd used to create the videos and ones in the later chunks may have been different algorithms from the start.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 738234,
          "author_name": "Yifan Xie",
          "author_url": "",
          "post_date": "2020-02-06T09:52:17.723000",
          "content": "<blockquote>\n  <p><strong>James Howard wrote:</strong></p>\n  \n  <p>Limiting it to my best kernel using a single model:\n  Validation loss: 0.1887\n  Leaderboard loss: 0.39309</p>\n</blockquote>\n\n<p>looks like some wonderful engineering work has been put in place to stack things up to give you the additional boost on LB 💯 </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 738238,
          "author_name": "Akash",
          "author_url": "",
          "post_date": "2020-02-06T09:56:38.003000",
          "content": "<p><a href=\"/jamesphoward\">@jamesphoward</a>  If you don't mind me asking, what resolution of images are you working with?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 738240,
          "author_name": "James Howard",
          "author_url": "",
          "post_date": "2020-02-06T10:01:57.917000",
          "content": "<p>I resize bits of interest to 256 x 256, with 0padding to make square, and then center-cropp to 224 x 224.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 772161,
          "author_name": "Alexey Kachalov",
          "author_url": "",
          "post_date": "2020-03-15T05:42:36.380000",
          "content": "<p><a href=\"/jamesphoward\">@jamesphoward</a>  What are do approximately? input for CNN is single face? Custom or public architecture?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 713013,
      "author_name": "Carlos Souza",
      "author_url": "",
      "post_date": "2020-01-07T20:14:40.837000",
      "content": "<p><a href=\"/harangdev\">@harangdev</a> , I did something similar, with the following tweaks:\n- 1/3 of the videos I kept unchanged\n- 2/9 of the videos I resized to 1/4 of their sizes\n- 2/9 of the videos I reduced FPS to 15\n- 2/9 of the videos I applied a hard compression</p>\n\n<p>This is exactly what is described in DFDC paper.</p>\n\n<p>My public LB is about the same as yours (~0.63), but my validation logloss is ~0.55, much closer to my LB.\nI suspect the key is the last bullet: apply a hard compression. This reduces the videos' file sizes to &lt;1/10 of their original sizes, and make it much harder for our algos to correctly classify as fake or real.\nIMPORTANT: I made sure these 4 proportions are respected in both training and validation sets.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 713244,
          "author_name": "YoonSoo",
          "author_url": "",
          "post_date": "2020-01-08T04:20:31.273000",
          "content": "<p>Yeah, I will implement compression augmentation later, but I'm not sure it will have impact on my validation score that much, like from 0.2 to 0.6. Thanks for your suggestions!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 714298,
          "author_name": "xxn-xx",
          "author_url": "",
          "post_date": "2020-01-09T09:26:09.903000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 733394,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-01-31T04:14:21.730000",
          "content": "<p>What exactly do you mean by hard compression? in albumentations we have an option for jpeg compression. Is that similar to what you used?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 733407,
          "author_name": "Ankit Saini",
          "author_url": "",
          "post_date": "2020-01-31T04:50:17.293000",
          "content": "<p>I'm also searching for the same. until now the best I can think to reduce the overall quality is to resize image size to (70,70) and then resize it to (224,224). Of course, numbers are just for illustration.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 733496,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-01-31T08:05:40.923000",
          "content": "<p>You can use both JpegCompression and Downscale from Albumentations:</p>\n\n<p><code>\nA.OneOf([\n    A.JpegCompression(quality_lower=8, quality_upper=30, p=1.0),\n    A.Downscale(scale_min=0.25, scale_max=0.75, p=1.0)\n ], p=0.22)\n</code></p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 742914,
          "author_name": "Cato",
          "author_url": "",
          "post_date": "2020-02-11T16:34:05.430000",
          "content": "<p><a href=\"/mpware\">@mpware</a> Why did you choose these Jpeg quality values? Are they related to the H.264 CRF values (e.g. default 23 corresponds to quality_upper=30)?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 743040,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-02-11T18:27:08.587000",
          "content": "<p>No, I've tried different settings to make JPEG compression hard.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 772401,
      "author_name": "GreatGameDota",
      "author_url": "",
      "post_date": "2020-03-15T12:53:26.560000",
      "content": "<p>I believe just splitting the folders 0-40, 41-49 gives a good validation method here's my results:\n```\nValidation Log Loss: .3347\nLB: .470</p>\n\n<p>Validation Log Loss: .3035\nLB: .446</p>\n\n<p>Validation Log Loss: .32555\nLB: .439\n<code>\nUPDATE:\nWith 0-39, 40-49 folder validation:\n</code>\nVal: .259\nLB: .392\n```\nThe trick is to save a checkpoint when the model is just about to overfit which is difficult. At best you can save every checkpoint and then only submit models that give a even 50/50 split on the test 400 vids.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 772421,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-03-15T13:20:56.687000",
          "content": "<p>My results were 0.160 on the same data, 0.318 on leaderboard.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 774974,
          "author_name": "Limerobot",
          "author_url": "",
          "post_date": "2020-03-16T06:05:38.137000",
          "content": "<p>Hi. I'm late joiner. <a href=\"/greatgamedota\">@greatgamedota</a> and <a href=\"/harshitsheoran\">@harshitsheoran</a> comments are really helpful. Thx!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 775166,
          "author_name": "Chason",
          "author_url": "",
          "post_date": "2020-03-16T11:24:02.767000",
          "content": "<p><a href=\"/harshitsheoran\">@harshitsheoran</a> Could you please share how you construct the validation set? I've tried a lot of strategies, but there's always a big gap between CV and LB.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 775213,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-03-16T12:18:46.373000",
          "content": "<p>Well, <a href=\"/chenshen03\">@chenshen03</a> it looks to me, that my validation gap is the largest, its double score  on leaderboard every time.</p>\n\n<p>Anyways, first list 10 last folders videos in a format of filenames, originals, then randomly take 500 of those, which are from my experience, testing was same for me as testing all folders all videos or just 500, which are chosen by shuffle([filenames,original], random_state=42)[:500] where shuffle is sklearn.utils.shuffle.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 712174,
      "author_name": "Shangqiu Li",
      "author_url": "",
      "post_date": "2020-01-06T22:47:08.167000",
      "content": "<p>My cv was 0.667 and my LB was 0.69243.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 712182,
          "author_name": "YoonSoo",
          "author_url": "",
          "post_date": "2020-01-06T23:03:54.747000",
          "content": "<p>Thank you for sharing. Your scores align pretty well. How did you split validation set?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 712202,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-01-07T00:00:45.207000",
          "content": "<p>I just found out a glitch in my validation strategy. I passed the string labels instead of number labels. The new CV is 0.69312.\nI split the validation set by folders(last two folders are the validation set).</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 712213,
          "author_name": "YoonSoo",
          "author_url": "",
          "post_date": "2020-01-07T00:23:14.360000",
          "content": "<p>Thanks for your reply. I think I should investigate further into my data splits.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 712217,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-01-07T00:30:18.297000",
          "content": "<p>Can you post the code here as well(if you don't feel like making the whole notebook public, just publish the key parts or replace the parts you want to keep secret with a brief summary of that part)? Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 712239,
          "author_name": "YoonSoo",
          "author_url": "",
          "post_date": "2020-01-07T01:26:01.580000",
          "content": "<p>```</p>\n\n<h1>load metadata to variable meta</h1>\n\n<p>meta['label'] = meta['label'].map({'FAKE': 1, 'REAL': 0})\nmeta = meta.set_index('filename')</p>\n\n<p>def split(meta, pos_val_size=2000):</p>\n\n<pre><code>meta = meta.sample(frac=1, random_state=0)\nsources = meta['original'].dropna().unique()\nvalid_pos = sources[:pos_val_size]\ntmp = meta.drop(valid_pos).reset_index().set_index('original')\nvalid_neg = tmp.loc[valid_pos, 'index'].groupby(level=0, group_keys=False).apply(lambda x: x.sample(1, random_state=0)).values\nvalid = np.concatenate([valid_pos, valid_neg])\ntrain_pos = sources[pos_val_size:]\ntrain_neg = tmp.loc[sources[train_pos], 'index'].values\ntrain = np.concatenate([train_pos, train_neg])\n\nnp.random.seed(0)\nnp.random.shuffle(train)\nnp.random.seed(0)\nnp.random.shuffle(valid)\ntrain = [x[:-4] for x in train]\nvalid = [x[:-4] for x in valid]\n\nreturn train, valid\n</code></pre>\n\n<p>train, valid = split(meta)\n```</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 714028,
      "author_name": "Pedro Bernardo",
      "author_url": "",
      "post_date": "2020-01-09T00:28:14.537000",
      "content": "<p>I've got cv:0.49 lb:0.53</p>\n\n<p>Apparently my strategy is working well enough as before I had cv:0.55/lb:0.59, so its tracking well.\nI'm using the strategy to split by file and downsample train and validation so it has roughly 50/50 fake and real.\nStill not using all the data available and doing a similar augmentation as the one <a href=\"/carlossouza\">@carlossouza</a> suggested here.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 714036,
          "author_name": "YoonSoo",
          "author_url": "",
          "post_date": "2020-01-09T00:58:36.507000",
          "content": "<p>Your scores align pretty well! I'm thinking towards that my local code has some bug in it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 724330,
          "author_name": "xxn-xx",
          "author_url": "",
          "post_date": "2020-01-21T04:00:48.180000",
          "content": "<p><a href=\"/pedromb\">@pedromb</a>  May I ask how you organize your validation set?  For example, fetch part 1,2,3,4,5 as validation or something like that?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 724622,
          "author_name": "Pedro Bernardo",
          "author_url": "",
          "post_date": "2020-01-21T10:03:39.230000",
          "content": "<p>I create a dataframe with the metadata and include a column that specify the part, then I use sklearn <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GroupShuffleSplit.html\">GroupShuffleSplit</a>.</p>\n\n<p>I just set the train size in a way that my test size has at least 2k real videos.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 735365,
          "author_name": "gody7334",
          "author_url": "",
          "post_date": "2020-02-02T23:36:10.333000",
          "content": "<p><a href=\"/pedromb\">@pedromb</a> do you think that the test dataset is 50/50 fake/real, so you split the same on validation set? \nin Train set, i use weighted sampling 50/50 get CV: 0.4\nI keep the original ratio 10/50 real/fake in validation set, get much better CV result: 0.2\nDoes it infer that test set distrubtion is closet to 50/50\nThanks</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 735395,
          "author_name": "pete",
          "author_url": "",
          "post_date": "2020-02-03T01:12:29.223000",
          "content": "<p>The test set is definitely 50:50.</p>\n\n<p><a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/126524\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/126524</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 735496,
      "author_name": "Shangqiu Li",
      "author_url": "",
      "post_date": "2020-02-03T05:35:19.333000",
      "content": "<p>EDIT:\ncv:0.0978 lb:0.0.40492</p>",
      "votes": 2,
      "replies": [
        {
          "id": 735526,
          "author_name": "Ankit Saini",
          "author_url": "",
          "post_date": "2020-02-03T06:43:34.913000",
          "content": "<p>You're doing great at cv. Seems like there are more audio manipulated videos in lb dataset. Can I know how you made the test-val split?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 736205,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-02-04T00:15:27.503000",
          "content": "<p>0.15 test_rate simple sklearn train_test_split. I will indeed try clustering for more accurate validation score.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 736426,
          "author_name": "Ankit Saini",
          "author_url": "",
          "post_date": "2020-02-04T07:01:04.580000",
          "content": "<p>I'm using first and last chuncks for validation (cv0.43, lb0.47). I think this helps due to variations in first set and audio manipulated videos in last set. Please share the results after clustering. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 738107,
          "author_name": "Chason",
          "author_url": "",
          "post_date": "2020-02-06T06:21:30.080000",
          "content": "<p><a href=\"/unkownhihi\">@unkownhihi</a> Is your CV loss based on image or video?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 733319,
      "author_name": "Rafi Hai",
      "author_url": "",
      "post_date": "2020-01-31T00:35:44.837000",
      "content": "<p><a href=\"/harangdev\">@harangdev</a>  I am wondering, what was the fix for your validation logloss? Was it the augmentations? \nI am also experiencing the discrepancy between the two.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 733321,
          "author_name": "YoonSoo",
          "author_url": "",
          "post_date": "2020-01-31T00:54:06.947000",
          "content": "<p>I haven't figure out how to make the cv and lb align yet.☹️</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 714324,
      "author_name": "FrazierLei",
      "author_url": "",
      "post_date": "2020-01-09T09:44:21.470000",
      "content": "<p>cv:0.48 lb:0.50, which looks normal</p>",
      "votes": 2,
      "replies": [
        {
          "id": 714443,
          "author_name": "YoonSoo",
          "author_url": "",
          "post_date": "2020-01-09T12:52:59.600000",
          "content": "<p>That's nice. Did you use full data?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 714477,
          "author_name": "FrazierLei",
          "author_url": "",
          "post_date": "2020-01-09T13:30:18.140000",
          "content": "<p>Almost the full data, I split apart some videos to try other methods.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 715009,
          "author_name": "YoonSoo",
          "author_url": "",
          "post_date": "2020-01-10T02:57:42.163000",
          "content": "<p><a href=\"/feifeizaici\">@feifeizaici</a> May I ask you how you split validation data? I'm still getting too good validation scores even after trying dozen experiments.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 715012,
          "author_name": "FrazierLei",
          "author_url": "",
          "post_date": "2020-01-10T03:10:59.197000",
          "content": "<p>I colleced 2000 FAKE and 2000 REAL videos, and for each of them I extracted 5 frames. My validation set is right these 20k faces. I think val/test images and train images should not come from a same batch of videos, for some frames are just too similar. </p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 715970,
          "author_name": "Chason",
          "author_url": "",
          "post_date": "2020-01-11T04:13:58.077000",
          "content": "<p><a href=\"/feifeizaici\">@feifeizaici</a> cv: 0.31, lb: 0.75. My model is overfitting. Do you have any data enhancement techniques? \nIn addition, the longer the training, the more severe the overfitting. To what extent does your model generally train? Thank you very much!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 715974,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-01-11T04:20:09.040000",
          "content": "<p>Looks like your score is a lot better than 0.75 😄 </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 715978,
          "author_name": "Chason",
          "author_url": "",
          "post_date": "2020-01-11T04:48:31.607000",
          "content": "<p><a href=\"/unkownhihi\">@unkownhihi</a>  My current score is obtained by randomly sampling 10K FAKE and REAL videos. Later, I tried to divide the training set and validation set according to origin videos, and increased the margin of face detection, but overfitting occurred.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 716051,
          "author_name": "FrazierLei",
          "author_url": "",
          "post_date": "2020-01-11T07:19:38.733000",
          "content": "<p>Just random horizontal flip. \nI think you can set a learning rate decay to aviod overfit.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 726548,
          "author_name": "pete",
          "author_url": "",
          "post_date": "2020-01-23T04:17:39.960000",
          "content": "<p><a href=\"/chenshen03\">@chenshen03</a> I seem to get the same thing happening when I use origin videos. Massive overfitting. Did you ever figure it out?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 713239,
      "author_name": "Harshit Sheoran",
      "author_url": "",
      "post_date": "2020-01-08T04:14:31.157000",
      "content": "<p>CV logloss is 0.08 using 8% of the total 500gb data for training and 2% for validating, I know this is not the greastest data to train my model on and I think that it is also overfitting on the faces of some people (confident to solve it) I mean validation accuracy is 96%, only a few tricks, but my score is stuck on 0.69314 does not matter how much I change the sub, I guess it is a bug in inference, If someone can help me solve it, I might go for making some of my code and tricks public sometime later...,\nAlso, if someone is super super good and confident, willing to work almost full time every day (I am saying that cuz last time I asked for teammates I got nobbies) can join my team after a test ofc. I only have a 1080 ti and a slow internet... (20 Mbps)</p>\n\n<p>I reken that the score is not gonna be good on the leaderboard the reason is the test data is a lot different then the data we are training on, I guess we will be able to solve that with some custom augs, I have no augs in my data currently...</p>\n\n<p><a href=\"/harangdev\">@harangdev</a> Thank you for this discussion, so many ideas I have not tried...</p>",
      "votes": 2,
      "replies": [
        {
          "id": 713247,
          "author_name": "YoonSoo",
          "author_url": "",
          "post_date": "2020-01-08T04:25:13.243000",
          "content": "<p>Hope you fix the bug in your inference code soon. I'm curious what your lb score will be.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 713253,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-01-08T04:35:34.013000",
          "content": "<p>Trust me this one is really gonna be bad I can feel it because which ever model I make I test it to some custom tests to get where the model is lacking which data it is not able to get right, And looking at the test set there are so many different people that if I dont train on full data it just does not make sense that model is gonna do any good then even random, also I dont have the pc to train it on full data, definately I can optimize it but, that will still take like 9 days for data creation and then 10-12 hours for each epoch that really pisses me off and I dont have the internet speed to just download the data I am downloading it every slowly ;-( basically that is the reason why I need a teammate...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 713331,
          "author_name": "Ankit Saini",
          "author_url": "",
          "post_date": "2020-01-08T06:57:17.587000",
          "content": "<p>You should go with google colab(with google drive) and transfer.sh I use combination of both. Transfer.sh is helpful for storing preprocessed videos and google drive for trained models. Then I have some wrote some functions to download a training set, uploading a folder to kaggle datatset, get face bounding box from a image as numpy array etc. I can provide if you're interested.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 713354,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-01-08T07:28:05.580000",
          "content": "<p>Why go with google colab???, it is slower than my pc in terms of processing power and I told you (copying this again lol) I can optimize it but, that will still take like 9 days for data creation and then 10-12 hours for each epoch that really pisses me off..., And god knows how much time on google colab...</p>\n\n<p>I have got a model for face detection which is much better than MTCNN and much faster too, the best I could find..., If you can answer why then sure I will be interested.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 713495,
          "author_name": "FrazierLei",
          "author_url": "",
          "post_date": "2020-01-08T10:52:21.213000",
          "content": "<p>Have you ever ploted your label's distribution in the kaggle kernel? Dose it looks normal?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 713497,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-01-08T10:54:26.140000",
          "content": "<p><a href=\"/feifeizaici\">@feifeizaici</a> I mean I have read the submission file kernel generated, the values looks as they should be to me</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 713829,
          "author_name": "Ankit Saini",
          "author_url": "",
          "post_date": "2020-01-08T17:34:24.800000",
          "content": "<p>I am not looking to join as team. I'll just help anyone having hard time to train to local setup.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 713959,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-08T21:13:44.250000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 714225,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-01-09T07:52:18.167000",
          "content": "<p>I guess I solved the problem which was making the model overfit on faces, now just 10 epoch training on the sample 400 videos, I am getting 0.4 loss on the dfdc_train_part_0 as the test set..., Someone help to solve the always same score bug. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 779819,
          "author_name": "A S M Iftekhar",
          "author_url": "",
          "post_date": "2020-03-19T17:52:14.727000",
          "content": "<p>May I know, what was the same score bug?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 712269,
      "author_name": "Marcus Lin",
      "author_url": "",
      "post_date": "2020-01-07T02:44:12.307000",
      "content": "<p>cv: 0.426  LB:17.26938 </p>",
      "votes": 2,
      "replies": [
        {
          "id": 712277,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-01-07T02:57:12.570000",
          "content": "<p>LB of 17.26938 is not normal. It can be solved by infering not over the sample_sumbission videos but over the /testvideos/* files. refer to <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/123260\">link</a>\nBTW cv of 0.426 is pretty good.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 712284,
          "author_name": "Marcus Lin",
          "author_url": "",
          "post_date": "2020-01-07T03:03:53.870000",
          "content": "<p>Nice to know 0.426 is good. I am trying to fixe submission file, will update.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 712287,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-01-07T03:19:04.893000",
          "content": "<p>Yeah! The best CV I can get is 0.666(unless overfit, hopefully your not overfitting). Its really not convenient that cuda can't produce reproducible results. My CV was good, but when I submit the notebook, it will train again and got bad results.  </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 713297,
          "author_name": "Ankit Saini",
          "author_url": "",
          "post_date": "2020-01-08T06:02:15.013000",
          "content": "<p>I can't see why inference over test_videos/ helps. Could you please elaborate.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 713932,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-01-08T20:19:09.567000",
          "content": "<p>The sample submission won't change but the filenames used to calculate LB will change.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 711597,
      "author_name": "nosound",
      "author_url": "",
      "post_date": "2020-01-06T09:41:20.423000",
      "content": "<p>I think it is not possible to get to validation log-loss 0.2 without a bug, especially when balanced. Do you use full train data? What is your train loss?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 711608,
          "author_name": "YoonSoo",
          "author_url": "",
          "post_date": "2020-01-06T10:05:46.547000",
          "content": "<p>I use full train data and my train loss is about 0.077.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 713191,
          "author_name": "Carlos Souza",
          "author_url": "",
          "post_date": "2020-01-08T02:25:40.977000",
          "content": "<p>Full train data is extremely imbalanced, approx 5:1 fake to real. If you don’t treat it, your model might show great results in training, but won’t generalize well as it will have a strong bias to predict as fake, thus generating a lot of false positives. Basically there are 3 ways of dealing with class imbalance: undersampling, oversampling, or weighting the loss function. If you are using Pytorch, the following link might be useful: <a href=\"https://discuss.pytorch.org/t/dealing-with-imbalanced-datasets-in-pytorch/22596\">https://discuss.pytorch.org/t/dealing-with-imbalanced-datasets-in-pytorch/22596</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 713195,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-01-08T02:27:52.653000",
          "content": "<p><a href=\"/harangdev\">@harangdev</a>  already said:</p>\n\n<blockquote>\n  <p>Make positive-negative ratio to 0.5 : 0.5 in validation set\n  So it is balanced.</p>\n</blockquote>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 711540,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-06T07:59:34.313000",
      "content": "<p>Seems interesting challenge... I'll also try this problem in few days. \nThanks for sharing your Views with us.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 748515,
      "author_name": "zjiang",
      "author_url": "",
      "post_date": "2020-02-17T16:07:28.673000",
      "content": "<p><a href=\"/harangdev\">@harangdev</a> Do you mind sharing your folder split number as Val set? Thanks</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 747364,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2020-02-16T10:41:54.490000",
      "content": "<p>Some great insights here. Could you share the source of your bug, might be useful? Thanks. :) </p>",
      "votes": 0,
      "replies": [
        {
          "id": 747374,
          "author_name": "YoonSoo",
          "author_url": "",
          "post_date": "2020-02-16T11:02:07.550000",
          "content": "<p>I wasn't familiar with model.train() and model.eval() thing. Also I moved from catalyst to vanilla pytorch.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 747408,
          "author_name": "Yassine Alouini",
          "author_url": "",
          "post_date": "2020-02-16T11:49:10.023000",
          "content": "<p>Thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 735488,
      "author_name": "Akash",
      "author_url": "",
      "post_date": "2020-02-03T05:21:04.673000",
      "content": "<p><a href=\"/harangdev\">@harangdev</a> \nHey, if you don't mind sharing, what changes did you make to your validation splits/pipeline in order to reduce LB from 0.6 to 0.39? Looks like I'm having a very similar issue. Was it a face overfitting that was happening for you? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 738191,
          "author_name": "YoonSoo",
          "author_url": "",
          "post_date": "2020-02-06T08:49:12.617000",
          "content": "<p>I've updated my post :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "711501": "These are validation strategies I've employed.\n\n1. Split train set and validation set with respect to source videos (see [here](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/124127))\n2. Make positive-negative ratio to 0.5 : 0.5 in validation set\n3. Simulated 1/4 resolution augmentations to 2/9 of the validation videos (see [here](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122013))\n\nAnd these are the scores I got.\n\n* validation logloss: around 0.20\n* public leaderboard logloss: around 0.60\n\nI'm suspecting three causes why the scores have such a large discrepancy.\n\n1. My inference notebook has some bug in it. (I've checked it thoroughly, but it is possible.)\n2. As discussed [here](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/124417), there may exist duplicates in train data, and I probably got duplicate videos both in train set and validation set.\n3. Maybe test set is comprised only of actors that don't appear in provided train set, and my model overfitted extremely to certain face identities.\n\nWhat is your validation score and public leaderboard score? (without leakage) Also, I'd appreciate it if you guys share your thoughts about this gap between scores.\n\n*Update) I fixed some bug in my code (I was new to pytorch and I had some issues), added some augmentations, etc. Just using folder split, not using augmentations for validation set. -&gt; Now validation logloss 0.22770, public logloss 0.39295*",
    "738174": "Limiting it to my best kernel using a single model:\nValidation loss: 0.1887\nLeaderboard loss: 0.39309",
    "713013": "@harangdev , I did something similar, with the following tweaks:\n- 1/3 of the videos I kept unchanged\n- 2/9 of the videos I resized to 1/4 of their sizes\n- 2/9 of the videos I reduced FPS to 15\n- 2/9 of the videos I applied a hard compression\n\nThis is exactly what is described in DFDC paper.\n\nMy public LB is about the same as yours (~0.63), but my validation logloss is ~0.55, much closer to my LB.\nI suspect the key is the last bullet: apply a hard compression. This reduces the videos' file sizes to &lt;1/10 of their original sizes, and make it much harder for our algos to correctly classify as fake or real.\nIMPORTANT: I made sure these 4 proportions are respected in both training and validation sets.",
    "772401": "I believe just splitting the folders 0-40, 41-49 gives a good validation method here's my results:\n```\nValidation Log Loss: .3347\nLB: .470\n\nValidation Log Loss: .3035\nLB: .446\n\nValidation Log Loss: .32555\nLB: .439\n```\nUPDATE:\nWith 0-39, 40-49 folder validation:\n```\nVal: .259\nLB: .392\n```\nThe trick is to save a checkpoint when the model is just about to overfit which is difficult. At best you can save every checkpoint and then only submit models that give a even 50/50 split on the test 400 vids.",
    "712174": "My cv was 0.667 and my LB was 0.69243.",
    "714028": "I've got cv:0.49 lb:0.53\n\nApparently my strategy is working well enough as before I had cv:0.55/lb:0.59, so its tracking well.\nI'm using the strategy to split by file and downsample train and validation so it has roughly 50/50 fake and real.\nStill not using all the data available and doing a similar augmentation as the one @carlossouza suggested here.",
    "735496": "EDIT:\ncv:0.0978 lb:0.0.40492",
    "733319": "@harangdev  I am wondering, what was the fix for your validation logloss? Was it the augmentations? \nI am also experiencing the discrepancy between the two.",
    "714324": "cv:0.48 lb:0.50, which looks normal\n\n",
    "713239": "CV logloss is 0.08 using 8% of the total 500gb data for training and 2% for validating, I know this is not the greastest data to train my model on and I think that it is also overfitting on the faces of some people (confident to solve it) I mean validation accuracy is 96%, only a few tricks, but my score is stuck on 0.69314 does not matter how much I change the sub, I guess it is a bug in inference, If someone can help me solve it, I might go for making some of my code and tricks public sometime later...,\nAlso, if someone is super super good and confident, willing to work almost full time every day (I am saying that cuz last time I asked for teammates I got nobbies) can join my team after a test ofc. I only have a 1080 ti and a slow internet... (20 Mbps)\n\nI reken that the score is not gonna be good on the leaderboard the reason is the test data is a lot different then the data we are training on, I guess we will be able to solve that with some custom augs, I have no augs in my data currently...\n\n@harangdev Thank you for this discussion, so many ideas I have not tried...",
    "712269": "cv: 0.426  LB:17.26938 ",
    "711597": "I think it is not possible to get to validation log-loss 0.2 without a bug, especially when balanced. Do you use full train data? What is your train loss?",
    "711540": "Seems interesting challenge... I'll also try this problem in few days. \nThanks for sharing your Views with us.",
    "748515": "@harangdev Do you mind sharing your folder split number as Val set? Thanks",
    "747364": "Some great insights here. Could you share the source of your bug, might be useful? Thanks. :) ",
    "735488": "@harangdev \nHey, if you don't mind sharing, what changes did you make to your validation splits/pipeline in order to reduce LB from 0.6 to 0.39? Looks like I'm having a very similar issue. Was it a face overfitting that was happening for you? "
  }
}