{
  "id": 122282,
  "title": "submission score is wrong ? ",
  "url": "/competitions/deepfake-detection-challenge/discussion/122282",
  "author_name": "",
  "post_date": "2019-12-19T03:15:23.570582Z",
  "votes": 4,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Hi guys, </p>\n\n<p>I submitted an output result, got log loss score <strong>17.26938</strong></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3270734%2Fce3c5f6d15cf195a49230f16687b7bf0%2F.PNG?generation=1576724716348647&amp;alt=media\" alt=\"\"></p>\n\n<p>for my submission, you can see screenshot, the results are okay, <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3270734%2Ff0a6dee708e2f6f46e41c6cb91449cd1%2F.PNG?generation=1576724995752281&amp;alt=media\" alt=\"\"></p>\n\n<p>and according to my calculation, my submission should be  <strong>0.8848960465390366</strong>.</p>\n\n<p>Besides if all prediction are zeros or ones, got </p>\n\n<p>log_loss([1, 0, 0, 1],[0, 0, 0, 0]) =  17.269388197455342 </p>\n\n<p>log_loss([1, 0, 0, 1],[1, 1, 1, 1]) =  17.269388197455342 </p>\n\n<p>so what's wrong with the submission scoring system?  anyone can help it. thanks a lot. </p>",
  "messages": [
    {
      "id": "698295",
      "postDate": "12/19/2019 03:15:23",
      "content": "<p>Hi guys, </p>\n\n<p>I submitted an output result, got log loss score <strong>17.26938</strong></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3270734%2Fce3c5f6d15cf195a49230f16687b7bf0%2F.PNG?generation=1576724716348647&amp;alt=media\" alt=\"\"></p>\n\n<p>for my submission, you can see screenshot, the results are okay, <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3270734%2Ff0a6dee708e2f6f46e41c6cb91449cd1%2F.PNG?generation=1576724995752281&amp;alt=media\" alt=\"\"></p>\n\n<p>and according to my calculation, my submission should be  <strong>0.8848960465390366</strong>.</p>\n\n<p>Besides if all prediction are zeros or ones, got </p>\n\n<p>log_loss([1, 0, 0, 1],[0, 0, 0, 0]) =  17.269388197455342 </p>\n\n<p>log_loss([1, 0, 0, 1],[1, 1, 1, 1]) =  17.269388197455342 </p>\n\n<p>so what's wrong with the submission scoring system?  anyone can help it. thanks a lot. </p>",
      "rawMarkdown": "Hi guys, \n\nI submitted an output result, got log loss score **17.26938**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3270734%2Fce3c5f6d15cf195a49230f16687b7bf0%2F.PNG?generation=1576724716348647&amp;alt=media)\n\nfor my submission, you can see screenshot, the results are okay, ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3270734%2Ff0a6dee708e2f6f46e41c6cb91449cd1%2F.PNG?generation=1576724995752281&amp;alt=media)\n\nand according to my calculation, my submission should be  **0.8848960465390366**.\n\nBesides if all prediction are zeros or ones, got \n\nlog_loss([1, 0, 0, 1],[0, 0, 0, 0]) =  17.269388197455342 \n\nlog_loss([1, 0, 0, 1],[1, 1, 1, 1]) =  17.269388197455342 \n\n\nso what's wrong with the submission scoring system?  anyone can help it. thanks a lot.",
      "votes": null
    },
    {
      "id": "698298",
      "postDate": "12/19/2019 03:23:37",
      "content": "<p>The submissions are being graded off of a different set than the one you are viewing here. Also, they may be using a different eps than you are using here. That information has not been made public. </p>",
      "rawMarkdown": "The submissions are being graded off of a different set than the one you are viewing here. Also, they may be using a different eps than you are using here. That information has not been made public.",
      "votes": null
    },
    {
      "id": "698299",
      "postDate": "12/19/2019 03:24:26",
      "content": "<p>Your submission may be erroring out and predicting all zeros or 1's as well. Hard to say without visibility</p>",
      "rawMarkdown": "Your submission may be erroring out and predicting all zeros or 1's as well. Hard to say without visibility",
      "votes": null
    },
    {
      "id": "698302",
      "postDate": "12/19/2019 03:32:40",
      "content": "<p>They use eps=1e-15 according to my calculations.\nEdit: based on this thread <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121348\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121348</a></p>\n\n<p>ln(1/sqrt(1e-15(1-1e-15)) = 17.26938</p>",
      "rawMarkdown": "They use eps=1e-15 according to my calculations.\nEdit: based on this thread https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121348\n\nln(1/sqrt(1e-15(1-1e-15)) = 17.26938",
      "votes": null
    },
    {
      "id": "698305",
      "postDate": "12/19/2019 03:41:08",
      "content": "<p>Yea that's probably it.</p>",
      "rawMarkdown": "Yea that's probably it.",
      "votes": null
    },
    {
      "id": "698310",
      "postDate": "12/19/2019 03:57:25",
      "content": "<p><a href=\"/ryches\">@ryches</a> <a href=\"/sdoria\">@sdoria</a>  thanks bros for your kind suggestions. I would check my notebook again and see where code was going wrongly.  </p>",
      "rawMarkdown": "ryches @sdoria  thanks bros for your kind suggestions. I would check my notebook again and see where code was going wrongly.",
      "votes": null
    },
    {
      "id": "698315",
      "postDate": "12/19/2019 04:06:46",
      "content": "<p><a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/overview/getting-started\">https://www.kaggle.com/c/deepfake-detection-challenge/overview/getting-started</a></p>\n\n<p>Public Test Set: This dataset is completely withheld and is what Kaggle’s platform computes the public leaderboard against. When you “Submit to Competition” from the “Output” file of a committed notebook that contains the competition’s dataset, your code will be re-run in the background against this Public Test Set. When the re-run is complete, the score will be posted to the public leaderboard. If the re-run fails, you will see an error reflected in your “My Submissions” page. Unfortunately, we are unable to surface any details about your error, so as to prevent error-probing. You are limited to 2 submissions per day, including submissions which error.</p>",
      "rawMarkdown": "https://www.kaggle.com/c/deepfake-detection-challenge/overview/getting-started\n\nPublic Test Set: This dataset is completely withheld and is what Kaggle’s platform computes the public leaderboard against. When you “Submit to Competition” from the “Output” file of a committed notebook that contains the competition’s dataset, your code will be re-run in the background against this Public Test Set. When the re-run is complete, the score will be posted to the public leaderboard. If the re-run fails, you will see an error reflected in your “My Submissions” page. Unfortunately, we are unable to surface any details about your error, so as to prevent error-probing. You are limited to 2 submissions per day, including submissions which error.",
      "votes": null
    },
    {
      "id": "698458",
      "postDate": "12/19/2019 09:10:34",
      "content": "<p>modified my code and submit again. another error , <strong>Submission CSV Not Found</strong></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3270734%2Fe963d438a516c74f364054f61828306d%2F.PNG?generation=1576746630227834&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "modified my code and submit again. another error , **Submission CSV Not Found**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3270734%2Fe963d438a516c74f364054f61828306d%2F.PNG?generation=1576746630227834&amp;alt=media)",
      "votes": null
    },
    {
      "id": "702717",
      "postDate": "12/25/2019 04:02:00",
      "content": "<p>hello, walton. I have made several submissions and the scores are always 17.26938. My label value is type float16. And I have clip my label in the range of 0.0001 to 0.99999. Can you please tell me how you fix this problem ?\nThx\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F263239%2F0d2e4d09c47f806d09a32c0fc54d670e%2Fsubmisson.png?generation=1577246517681358&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "hello, walton. I have made several submissions and the scores are always 17.26938. My label value is type float16. And I have clip my label in the range of 0.0001 to 0.99999. Can you please tell me how you fix this problem ?\nThx\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F263239%2F0d2e4d09c47f806d09a32c0fc54d670e%2Fsubmisson.png?generation=1577246517681358&amp;alt=media)",
      "votes": null
    },
    {
      "id": "702871",
      "postDate": "12/25/2019 09:32:13",
      "content": "<p>You should not use that submission.csv. You should read the videos from <code>test_videos</code> and make predictions on those. They are different videos than are in the <code>sample_submission.csv</code> file.</p>",
      "rawMarkdown": "You should not use that submission.csv. You should read the videos from `test_videos` and make predictions on those. They are different videos than are in the `sample_submission.csv` file.",
      "votes": null
    },
    {
      "id": "703350",
      "postDate": "12/26/2019 03:17:56",
      "content": "<p>hi chien,</p>\n\n<p>so far I guess my code running a long time and caused the failure. because the private test dataset has aound  4000 videos. If cannot run this in 2 hours, it would cause the error. </p>",
      "rawMarkdown": "hi chien,\n\nso far I guess my code running a long time and caused the failure. because the private test dataset has aound  4000 videos. If cannot run this in 2 hours, it would cause the error.",
      "votes": null
    },
    {
      "id": "707601",
      "postDate": "01/01/2020 06:05:17",
      "content": "<p><a href=\"/humananalog\">@humananalog</a> We know that the <code>sample_submission.csv</code> file that we can see (public to us) contains 400 rows because the visible test set has 400 videos. When we submit the kernel for scoring, will the <code>sample_submission.csv</code> file now have 4000 rows as we know that there are approx 4000 hidden test videos or is this file the same as before (with only 400 rows)?\nI want to just map the output probability with the corresponding \"filename\" from the <code>sample_submission.csv</code> file instead of creating a new dataframe.\nI am asking this because I am also getting a score of <code>17.26938</code> but my predictions on visible test set are great and very accurate.</p>\n\n<p>Code I am using to write predictions to submissions.csv file (for both visible and invisible test set).\n```python\nfilelist = sorted(glob.glob('/kaggle/input/deepfake-detection-challenge/test_videos/*'), key=numericalSort)</p>\n\n<p>batch_size = 20\npreds = model.predict_generator(DataGenerator_predict(filelist=filelist), steps = np.ceil(len(filelist)/batch_size),verbose=1)\n```</p>\n\n<p><strong>What I am doing</strong>: Reading the <code>sample_submission.csv</code> file and editing it\n```python\nsub = pd.read_csv(\"/kaggle/input/deepfake-detection-challenge/sample_submission.csv\")</p>\n\n<p>sub.label = sub.label.astype(float)\nfor i,item in enumerate(filelist):\n    fname = item.split('/')[-1]\n    index = sub['filename'].tolist().index(fname)\n    if preds[i,1] &gt; 0.9:\n        sub['label'][index] = 0.9\n    elif preds[i,1] &gt;= 0.09 and preds[i,1] &lt;= 0.9:\n        sub['label'][index] = preds[i,1]\n    elif preds[i,1] &lt; 0.09:\n        sub['label'][index] = 0.09</p>\n\n<p>sub.fillna(0.5)</p>\n\n<p>sub.to_csv('submission.csv',index=False)\n```</p>\n\n<p>or should I do something like this without reading <code>sample_submission.csv</code>?</p>\n\n<p>```python\nout = []\nfor i,item in enumerate(filelist):\n    fname = item.split('/')[-1]\n    #index = sub['filename'].tolist().index(fname)\n    out1 = [fname, preds[i,1]]\n    out.append(out1)</p>\n\n<p>df = pd.DataFrame(out, columns=[\"filename\", \"label\"])</p>\n\n<p>df.to_csv('submission.csv',index=False)\n```</p>",
      "rawMarkdown": "humananalog We know that the `sample_submission.csv` file that we can see (public to us) contains 400 rows because the visible test set has 400 videos. When we submit the kernel for scoring, will the `sample_submission.csv` file now have 4000 rows as we know that there are approx 4000 hidden test videos or is this file the same as before (with only 400 rows)?\nI want to just map the output probability with the corresponding \"filename\" from the `sample_submission.csv` file instead of creating a new dataframe.\nI am asking this because I am also getting a score of `17.26938` but my predictions on visible test set are great and very accurate.\n\nCode I am using to write predictions to submissions.csv file (for both visible and invisible test set).\n```python\nfilelist = sorted(glob.glob('/kaggle/input/deepfake-detection-challenge/test_videos/*'), key=numericalSort)\n\nbatch_size = 20\npreds = model.predict_generator(DataGenerator_predict(filelist=filelist), steps = np.ceil(len(filelist)/batch_size),verbose=1)\n```\n\n**What I am doing**: Reading the `sample_submission.csv` file and editing it\n```python\nsub = pd.read_csv(\"/kaggle/input/deepfake-detection-challenge/sample_submission.csv\")\n\nsub.label = sub.label.astype(float)\nfor i,item in enumerate(filelist):\n    fname = item.split('/')[-1]\n    index = sub['filename'].tolist().index(fname)\n    if preds[i,1] &gt; 0.9:\n        sub['label'][index] = 0.9\n    elif preds[i,1] &gt;= 0.09 and preds[i,1] &lt;= 0.9:\n        sub['label'][index] = preds[i,1]\n    elif preds[i,1] &lt; 0.09:\n        sub['label'][index] = 0.09\n\nsub.fillna(0.5)\n\nsub.to_csv('submission.csv',index=False)\n```\n\nor should I do something like this without reading `sample_submission.csv`?\n\n```python\nout = []\nfor i,item in enumerate(filelist):\n    fname = item.split('/')[-1]\n    #index = sub['filename'].tolist().index(fname)\n    out1 = [fname, preds[i,1]]\n    out.append(out1)\n\ndf = pd.DataFrame(out, columns=[\"filename\", \"label\"])\n\ndf.to_csv('submission.csv',index=False)\n```",
      "votes": null
    },
    {
      "id": "707733",
      "postDate": "01/01/2020 11:24:08",
      "content": "<p>From what other people have said, yes I do think the <code>sample_submission.csv</code> also gets replaced at test time so you should be able to load that and fill it in. (So my earlier comment on this was wrong.)</p>\n\n<p>What I would try is don't call your model.predict stuff but choose 4000 random numbers and fill those in. Does your submission work then? If yes, something goes wrong with making the predictions.</p>",
      "rawMarkdown": "From what other people have said, yes I do think the `sample_submission.csv` also gets replaced at test time so you should be able to load that and fill it in. (So my earlier comment on this was wrong.)\n\nWhat I would try is don't call your model.predict stuff but choose 4000 random numbers and fill those in. Does your submission work then? If yes, something goes wrong with making the predictions.",
      "votes": null
    },
    {
      "id": "707821",
      "postDate": "01/01/2020 14:27:00",
      "content": "<p>Do we know that the public test set has 4000 values exactly?</p>",
      "rawMarkdown": "Do we know that the public test set has 4000 values exactly?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 698298,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "12/19/2019 03:23:37",
      "content": "<p>The submissions are being graded off of a different set than the one you are viewing here. Also, they may be using a different eps than you are using here. That information has not been made public. </p>",
      "votes": null,
      "replies": [
        {
          "id": 698302,
          "author_name": "sdoria",
          "author_url": "",
          "post_date": "12/19/2019 03:32:40",
          "content": "<p>They use eps=1e-15 according to my calculations.\nEdit: based on this thread <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121348\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121348</a></p>\n\n<p>ln(1/sqrt(1e-15(1-1e-15)) = 17.26938</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 698299,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "12/19/2019 03:24:26",
      "content": "<p>Your submission may be erroring out and predicting all zeros or 1's as well. Hard to say without visibility</p>",
      "votes": null,
      "replies": [
        {
          "id": 698305,
          "author_name": "sdoria",
          "author_url": "",
          "post_date": "12/19/2019 03:41:08",
          "content": "<p>Yea that's probably it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 698310,
      "author_name": "wangtao360",
      "author_url": "",
      "post_date": "12/19/2019 03:57:25",
      "content": "<p><a href=\"/ryches\">@ryches</a> <a href=\"/sdoria\">@sdoria</a>  thanks bros for your kind suggestions. I would check my notebook again and see where code was going wrongly.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 698315,
      "author_name": "wangtao360",
      "author_url": "",
      "post_date": "12/19/2019 04:06:46",
      "content": "<p><a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/overview/getting-started\">https://www.kaggle.com/c/deepfake-detection-challenge/overview/getting-started</a></p>\n\n<p>Public Test Set: This dataset is completely withheld and is what Kaggle’s platform computes the public leaderboard against. When you “Submit to Competition” from the “Output” file of a committed notebook that contains the competition’s dataset, your code will be re-run in the background against this Public Test Set. When the re-run is complete, the score will be posted to the public leaderboard. If the re-run fails, you will see an error reflected in your “My Submissions” page. Unfortunately, we are unable to surface any details about your error, so as to prevent error-probing. You are limited to 2 submissions per day, including submissions which error.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 698458,
      "author_name": "wangtao360",
      "author_url": "",
      "post_date": "12/19/2019 09:10:34",
      "content": "<p>modified my code and submit again. another error , <strong>Submission CSV Not Found</strong></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3270734%2Fe963d438a516c74f364054f61828306d%2F.PNG?generation=1576746630227834&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 702717,
          "author_name": "ericji",
          "author_url": "",
          "post_date": "12/25/2019 04:02:00",
          "content": "<p>hello, walton. I have made several submissions and the scores are always 17.26938. My label value is type float16. And I have clip my label in the range of 0.0001 to 0.99999. Can you please tell me how you fix this problem ?\nThx\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F263239%2F0d2e4d09c47f806d09a32c0fc54d670e%2Fsubmisson.png?generation=1577246517681358&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 702871,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "12/25/2019 09:32:13",
          "content": "<p>You should not use that submission.csv. You should read the videos from <code>test_videos</code> and make predictions on those. They are different videos than are in the <code>sample_submission.csv</code> file.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 703350,
          "author_name": "wangtao360",
          "author_url": "",
          "post_date": "12/26/2019 03:17:56",
          "content": "<p>hi chien,</p>\n\n<p>so far I guess my code running a long time and caused the failure. because the private test dataset has aound  4000 videos. If cannot run this in 2 hours, it would cause the error. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 707601,
          "author_name": "manideep2510",
          "author_url": "",
          "post_date": "01/01/2020 06:05:17",
          "content": "<p><a href=\"/humananalog\">@humananalog</a> We know that the <code>sample_submission.csv</code> file that we can see (public to us) contains 400 rows because the visible test set has 400 videos. When we submit the kernel for scoring, will the <code>sample_submission.csv</code> file now have 4000 rows as we know that there are approx 4000 hidden test videos or is this file the same as before (with only 400 rows)?\nI want to just map the output probability with the corresponding \"filename\" from the <code>sample_submission.csv</code> file instead of creating a new dataframe.\nI am asking this because I am also getting a score of <code>17.26938</code> but my predictions on visible test set are great and very accurate.</p>\n\n<p>Code I am using to write predictions to submissions.csv file (for both visible and invisible test set).\n```python\nfilelist = sorted(glob.glob('/kaggle/input/deepfake-detection-challenge/test_videos/*'), key=numericalSort)</p>\n\n<p>batch_size = 20\npreds = model.predict_generator(DataGenerator_predict(filelist=filelist), steps = np.ceil(len(filelist)/batch_size),verbose=1)\n```</p>\n\n<p><strong>What I am doing</strong>: Reading the <code>sample_submission.csv</code> file and editing it\n```python\nsub = pd.read_csv(\"/kaggle/input/deepfake-detection-challenge/sample_submission.csv\")</p>\n\n<p>sub.label = sub.label.astype(float)\nfor i,item in enumerate(filelist):\n    fname = item.split('/')[-1]\n    index = sub['filename'].tolist().index(fname)\n    if preds[i,1] &gt; 0.9:\n        sub['label'][index] = 0.9\n    elif preds[i,1] &gt;= 0.09 and preds[i,1] &lt;= 0.9:\n        sub['label'][index] = preds[i,1]\n    elif preds[i,1] &lt; 0.09:\n        sub['label'][index] = 0.09</p>\n\n<p>sub.fillna(0.5)</p>\n\n<p>sub.to_csv('submission.csv',index=False)\n```</p>\n\n<p>or should I do something like this without reading <code>sample_submission.csv</code>?</p>\n\n<p>```python\nout = []\nfor i,item in enumerate(filelist):\n    fname = item.split('/')[-1]\n    #index = sub['filename'].tolist().index(fname)\n    out1 = [fname, preds[i,1]]\n    out.append(out1)</p>\n\n<p>df = pd.DataFrame(out, columns=[\"filename\", \"label\"])</p>\n\n<p>df.to_csv('submission.csv',index=False)\n```</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 707733,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "01/01/2020 11:24:08",
          "content": "<p>From what other people have said, yes I do think the <code>sample_submission.csv</code> also gets replaced at test time so you should be able to load that and fill it in. (So my earlier comment on this was wrong.)</p>\n\n<p>What I would try is don't call your model.predict stuff but choose 4000 random numbers and fill those in. Does your submission work then? If yes, something goes wrong with making the predictions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 707821,
          "author_name": "sdoria",
          "author_url": "",
          "post_date": "01/01/2020 14:27:00",
          "content": "<p>Do we know that the public test set has 4000 values exactly?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "698295": "Hi guys, \n\nI submitted an output result, got log loss score **17.26938**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3270734%2Fce3c5f6d15cf195a49230f16687b7bf0%2F.PNG?generation=1576724716348647&amp;alt=media)\n\nfor my submission, you can see screenshot, the results are okay, ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3270734%2Ff0a6dee708e2f6f46e41c6cb91449cd1%2F.PNG?generation=1576724995752281&amp;alt=media)\n\nand according to my calculation, my submission should be  **0.8848960465390366**.\n\nBesides if all prediction are zeros or ones, got \n\nlog_loss([1, 0, 0, 1],[0, 0, 0, 0]) =  17.269388197455342 \n\nlog_loss([1, 0, 0, 1],[1, 1, 1, 1]) =  17.269388197455342 \n\n\nso what's wrong with the submission scoring system?  anyone can help it. thanks a lot.",
    "698298": "The submissions are being graded off of a different set than the one you are viewing here. Also, they may be using a different eps than you are using here. That information has not been made public.",
    "698299": "Your submission may be erroring out and predicting all zeros or 1's as well. Hard to say without visibility",
    "698302": "They use eps=1e-15 according to my calculations.\nEdit: based on this thread https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121348\n\nln(1/sqrt(1e-15(1-1e-15)) = 17.26938",
    "698305": "Yea that's probably it.",
    "698310": "ryches @sdoria  thanks bros for your kind suggestions. I would check my notebook again and see where code was going wrongly.",
    "698315": "https://www.kaggle.com/c/deepfake-detection-challenge/overview/getting-started\n\nPublic Test Set: This dataset is completely withheld and is what Kaggle’s platform computes the public leaderboard against. When you “Submit to Competition” from the “Output” file of a committed notebook that contains the competition’s dataset, your code will be re-run in the background against this Public Test Set. When the re-run is complete, the score will be posted to the public leaderboard. If the re-run fails, you will see an error reflected in your “My Submissions” page. Unfortunately, we are unable to surface any details about your error, so as to prevent error-probing. You are limited to 2 submissions per day, including submissions which error.",
    "698458": "modified my code and submit again. another error , **Submission CSV Not Found**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3270734%2Fe963d438a516c74f364054f61828306d%2F.PNG?generation=1576746630227834&amp;alt=media)",
    "702717": "hello, walton. I have made several submissions and the scores are always 17.26938. My label value is type float16. And I have clip my label in the range of 0.0001 to 0.99999. Can you please tell me how you fix this problem ?\nThx\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F263239%2F0d2e4d09c47f806d09a32c0fc54d670e%2Fsubmisson.png?generation=1577246517681358&amp;alt=media)",
    "702871": "You should not use that submission.csv. You should read the videos from `test_videos` and make predictions on those. They are different videos than are in the `sample_submission.csv` file.",
    "703350": "hi chien,\n\nso far I guess my code running a long time and caused the failure. because the private test dataset has aound  4000 videos. If cannot run this in 2 hours, it would cause the error.",
    "707601": "humananalog We know that the `sample_submission.csv` file that we can see (public to us) contains 400 rows because the visible test set has 400 videos. When we submit the kernel for scoring, will the `sample_submission.csv` file now have 4000 rows as we know that there are approx 4000 hidden test videos or is this file the same as before (with only 400 rows)?\nI want to just map the output probability with the corresponding \"filename\" from the `sample_submission.csv` file instead of creating a new dataframe.\nI am asking this because I am also getting a score of `17.26938` but my predictions on visible test set are great and very accurate.\n\nCode I am using to write predictions to submissions.csv file (for both visible and invisible test set).\n```python\nfilelist = sorted(glob.glob('/kaggle/input/deepfake-detection-challenge/test_videos/*'), key=numericalSort)\n\nbatch_size = 20\npreds = model.predict_generator(DataGenerator_predict(filelist=filelist), steps = np.ceil(len(filelist)/batch_size),verbose=1)\n```\n\n**What I am doing**: Reading the `sample_submission.csv` file and editing it\n```python\nsub = pd.read_csv(\"/kaggle/input/deepfake-detection-challenge/sample_submission.csv\")\n\nsub.label = sub.label.astype(float)\nfor i,item in enumerate(filelist):\n    fname = item.split('/')[-1]\n    index = sub['filename'].tolist().index(fname)\n    if preds[i,1] &gt; 0.9:\n        sub['label'][index] = 0.9\n    elif preds[i,1] &gt;= 0.09 and preds[i,1] &lt;= 0.9:\n        sub['label'][index] = preds[i,1]\n    elif preds[i,1] &lt; 0.09:\n        sub['label'][index] = 0.09\n\nsub.fillna(0.5)\n\nsub.to_csv('submission.csv',index=False)\n```\n\nor should I do something like this without reading `sample_submission.csv`?\n\n```python\nout = []\nfor i,item in enumerate(filelist):\n    fname = item.split('/')[-1]\n    #index = sub['filename'].tolist().index(fname)\n    out1 = [fname, preds[i,1]]\n    out.append(out1)\n\ndf = pd.DataFrame(out, columns=[\"filename\", \"label\"])\n\ndf.to_csv('submission.csv',index=False)\n```",
    "707733": "From what other people have said, yes I do think the `sample_submission.csv` also gets replaced at test time so you should be able to load that and fill it in. (So my earlier comment on this was wrong.)\n\nWhat I would try is don't call your model.predict stuff but choose 4000 random numbers and fill those in. Does your submission work then? If yes, something goes wrong with making the predictions.",
    "707821": "Do we know that the public test set has 4000 values exactly?"
  },
  "source": "meta"
}