{
  "id": 123260,
  "title": "Submission allways scores 0.69314",
  "url": "/competitions/deepfake-detection-challenge/discussion/123260",
  "author_name": "",
  "post_date": "2019-12-26T08:09:24.317485800Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Firstly I got <code>Submission CSV Not Found</code> that was solved by puting a df.head() before saving to csv. After that I got <code>17.26938</code> submissions that were solved by infering not over the sample__sumbission videos but over the /test_videos/* files. Now I got <code>0.69314</code> no matter how I modify the predictions. </p>\n\n<p>I'm asking you because I don't want to waste more submissions. Anyone has suffer the same issue? Could It be the df's label column dtype (I'm using float64)?</p>\n\n<p>Thank you in advance.</p>",
  "messages": [
    {
      "id": "703465",
      "postDate": "12/26/2019 08:09:24",
      "content": "<p>Firstly I got <code>Submission CSV Not Found</code> that was solved by puting a df.head() before saving to csv. After that I got <code>17.26938</code> submissions that were solved by infering not over the sample__sumbission videos but over the /test_videos/* files. Now I got <code>0.69314</code> no matter how I modify the predictions. </p>\n\n<p>I'm asking you because I don't want to waste more submissions. Anyone has suffer the same issue? Could It be the df's label column dtype (I'm using float64)?</p>\n\n<p>Thank you in advance.</p>",
      "rawMarkdown": "Firstly I got `Submission CSV Not Found` that was solved by puting a df.head() before saving to csv. After that I got `17.26938` submissions that were solved by infering not over the sample__sumbission videos but over the /test_videos/* files. Now I got `0.69314` no matter how I modify the predictions. \n\nI'm asking you because I don't want to waste more submissions. Anyone has suffer the same issue? Could It be the df's label column dtype (I'm using float64)?\n\nThank you in advance.",
      "votes": null
    },
    {
      "id": "704141",
      "postDate": "12/27/2019 05:05:05",
      "content": "<p>0.69314 seems to be -log(0.5), which is the expected score you'd get if every prediction were 0.5. If the guesses are coming from training, then it sounds like an overfitting problem where the guesses are converging to do just as good as random guessing would. If you're just testing to see if it's working properly, then what kind of inputs are you testing with? For instance, guessing randomly from a uniform distribution should yield a score of 1.0.</p>",
      "rawMarkdown": "0.69314 seems to be -log(0.5), which is the expected score you'd get if every prediction were 0.5. If the guesses are coming from training, then it sounds like an overfitting problem where the guesses are converging to do just as good as random guessing would. If you're just testing to see if it's working properly, then what kind of inputs are you testing with? For instance, guessing randomly from a uniform distribution should yield a score of 1.0.",
      "votes": null
    },
    {
      "id": "704316",
      "postDate": "12/27/2019 10:05:37",
      "content": "<p>I reckon it is not overfitting because I have performed CV over multiple parts of the dataset and the results are great.</p>",
      "rawMarkdown": "I reckon it is not overfitting because I have performed CV over multiple parts of the dataset and the results are great.",
      "votes": null
    },
    {
      "id": "713296",
      "postDate": "01/08/2020 05:56:32",
      "content": "<p>I'm also getting score around 15-17 and around 0.7(80% accuracy) on my validation (taken from completely different folder from training part40) .\nI can't understand why inference over test_videos rather then sample_submission helps. It will be helpful if you could explain.</p>",
      "rawMarkdown": "I'm also getting score around 15-17 and around 0.7(80% accuracy) on my validation (taken from completely different folder from training part40) .\nI can't understand why inference over test_videos rather then sample_submission helps. It will be helpful if you could explain.",
      "votes": null
    },
    {
      "id": "713305",
      "postDate": "01/08/2020 06:14:14",
      "content": "<p>that's my issue too, how to generate submission.csv filename sequence seems we can't use sample_submission.csv.</p>",
      "rawMarkdown": "that's my issue too, how to generate submission.csv filename sequence seems we can't use sample_submission.csv.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 704141,
      "author_name": "eeegnu",
      "author_url": "",
      "post_date": "12/27/2019 05:05:05",
      "content": "<p>0.69314 seems to be -log(0.5), which is the expected score you'd get if every prediction were 0.5. If the guesses are coming from training, then it sounds like an overfitting problem where the guesses are converging to do just as good as random guessing would. If you're just testing to see if it's working properly, then what kind of inputs are you testing with? For instance, guessing randomly from a uniform distribution should yield a score of 1.0.</p>",
      "votes": null,
      "replies": [
        {
          "id": 704316,
          "author_name": "cayala",
          "author_url": "",
          "post_date": "12/27/2019 10:05:37",
          "content": "<p>I reckon it is not overfitting because I have performed CV over multiple parts of the dataset and the results are great.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 713296,
      "author_name": "ankitsainiankit",
      "author_url": "",
      "post_date": "01/08/2020 05:56:32",
      "content": "<p>I'm also getting score around 15-17 and around 0.7(80% accuracy) on my validation (taken from completely different folder from training part40) .\nI can't understand why inference over test_videos rather then sample_submission helps. It will be helpful if you could explain.</p>",
      "votes": null,
      "replies": [
        {
          "id": 713305,
          "author_name": "marcuslin",
          "author_url": "",
          "post_date": "01/08/2020 06:14:14",
          "content": "<p>that's my issue too, how to generate submission.csv filename sequence seems we can't use sample_submission.csv.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "703465": "Firstly I got `Submission CSV Not Found` that was solved by puting a df.head() before saving to csv. After that I got `17.26938` submissions that were solved by infering not over the sample__sumbission videos but over the /test_videos/* files. Now I got `0.69314` no matter how I modify the predictions. \n\nI'm asking you because I don't want to waste more submissions. Anyone has suffer the same issue? Could It be the df's label column dtype (I'm using float64)?\n\nThank you in advance.",
    "704141": "0.69314 seems to be -log(0.5), which is the expected score you'd get if every prediction were 0.5. If the guesses are coming from training, then it sounds like an overfitting problem where the guesses are converging to do just as good as random guessing would. If you're just testing to see if it's working properly, then what kind of inputs are you testing with? For instance, guessing randomly from a uniform distribution should yield a score of 1.0.",
    "704316": "I reckon it is not overfitting because I have performed CV over multiple parts of the dataset and the results are great.",
    "713296": "I'm also getting score around 15-17 and around 0.7(80% accuracy) on my validation (taken from completely different folder from training part40) .\nI can't understand why inference over test_videos rather then sample_submission helps. It will be helpful if you could explain.",
    "713305": "that's my issue too, how to generate submission.csv filename sequence seems we can't use sample_submission.csv."
  },
  "source": "meta"
}