{
  "id": 127910,
  "title": "Leaderboard score based on test or train log loss?",
  "url": "/competitions/deepfake-detection-challenge/discussion/127910",
  "author_name": "",
  "post_date": "2020-01-27T15:54:19.058272100Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>hey Y'all!</p>\n\n<p>I'm considering submitting the inference of my trained model on test videos but I'm slightly confused here, any help is appreciated! \nQuestion: In order to calculate the log loss we are required to know the ground truth, which we are not provided with for test videos. How do I calculate log loss for test videos? </p>\n\n<p>Thanks in advance</p>",
  "messages": [
    {
      "id": "730540",
      "postDate": "01/27/2020 15:54:19",
      "content": "<p>hey Y'all!</p>\n\n<p>I'm considering submitting the inference of my trained model on test videos but I'm slightly confused here, any help is appreciated! \nQuestion: In order to calculate the log loss we are required to know the ground truth, which we are not provided with for test videos. How do I calculate log loss for test videos? </p>\n\n<p>Thanks in advance</p>",
      "rawMarkdown": "hey Y'all!\n\nI'm considering submitting the inference of my trained model on test videos but I'm slightly confused here, any help is appreciated! \nQuestion: In order to calculate the log loss we are required to know the ground truth, which we are not provided with for test videos. How do I calculate log loss for test videos? \n\nThanks in advance",
      "votes": null
    },
    {
      "id": "730553",
      "postDate": "01/27/2020 16:11:17",
      "content": "<p>Actually the test set is generated from parts of the full training dataset and you could potentially pull out the ground truth.</p>\n\n<p>Or instead you can separate a portion of your training set (ie 10%) and not train with it and then calculate log loss on that after you train your model.</p>",
      "rawMarkdown": "Actually the test set is generated from parts of the full training dataset and you could potentially pull out the ground truth.\n\nOr instead you can separate a portion of your training set (ie 10%) and not train with it and then calculate log loss on that after you train your model.",
      "votes": null
    },
    {
      "id": "730566",
      "postDate": "01/27/2020 16:25:43",
      "content": "<p>Gotcha. That seems doable, but wouldn't that mean all participants have different ways of calculating log loss? On different parts of the data, I mean?</p>",
      "rawMarkdown": "Gotcha. That seems doable, but wouldn't that mean all participants have different ways of calculating log loss? On different parts of the data, I mean?",
      "votes": null
    },
    {
      "id": "730597",
      "postDate": "01/27/2020 17:06:41",
      "content": "<p>You need to build a model to predict the probabilities of each sample being a deep fake, this will be a value between 0 and 1, 1.0 being very sure it's a fake, 0 being very sure it's real, 0.5 would mean your model is unsure. Kaggle know the labels and will calculate the log loss of your model's predictions when you submit your solution. </p>",
      "rawMarkdown": "You need to build a model to predict the probabilities of each sample being a deep fake, this will be a value between 0 and 1, 1.0 being very sure it's a fake, 0 being very sure it's real, 0.5 would mean your model is unsure. Kaggle know the labels and will calculate the log loss of your model's predictions when you submit your solution.",
      "votes": null
    },
    {
      "id": "730651",
      "postDate": "01/27/2020 18:25:16",
      "content": "<p>So you include your model in the inference notebook. When you submit the notebook, the test video dir will be swapped to an unseen validation set. Then, after your model predicts that unseen validation set, it should output as submission.csv. Then kaggle will calculate log loss based on your prediction(submission.csv)</p>",
      "rawMarkdown": "So you include your model in the inference notebook. When you submit the notebook, the test video dir will be swapped to an unseen validation set. Then, after your model predicts that unseen validation set, it should output as submission.csv. Then kaggle will calculate log loss based on your prediction(submission.csv)",
      "votes": null
    },
    {
      "id": "730714",
      "postDate": "01/27/2020 20:07:05",
      "content": "<p>Perfect! Thank you for the explanations.</p>",
      "rawMarkdown": "Perfect! Thank you for the explanations.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 730553,
      "author_name": "sethkitchen",
      "author_url": "",
      "post_date": "01/27/2020 16:11:17",
      "content": "<p>Actually the test set is generated from parts of the full training dataset and you could potentially pull out the ground truth.</p>\n\n<p>Or instead you can separate a portion of your training set (ie 10%) and not train with it and then calculate log loss on that after you train your model.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 730566,
      "author_name": "shivangeetrivedi",
      "author_url": "",
      "post_date": "01/27/2020 16:25:43",
      "content": "<p>Gotcha. That seems doable, but wouldn't that mean all participants have different ways of calculating log loss? On different parts of the data, I mean?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 730597,
      "author_name": "jackvial",
      "author_url": "",
      "post_date": "01/27/2020 17:06:41",
      "content": "<p>You need to build a model to predict the probabilities of each sample being a deep fake, this will be a value between 0 and 1, 1.0 being very sure it's a fake, 0 being very sure it's real, 0.5 would mean your model is unsure. Kaggle know the labels and will calculate the log loss of your model's predictions when you submit your solution. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 730651,
      "author_name": "unkownhihi",
      "author_url": "",
      "post_date": "01/27/2020 18:25:16",
      "content": "<p>So you include your model in the inference notebook. When you submit the notebook, the test video dir will be swapped to an unseen validation set. Then, after your model predicts that unseen validation set, it should output as submission.csv. Then kaggle will calculate log loss based on your prediction(submission.csv)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 730714,
      "author_name": "shivangeetrivedi",
      "author_url": "",
      "post_date": "01/27/2020 20:07:05",
      "content": "<p>Perfect! Thank you for the explanations.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "730540": "hey Y'all!\n\nI'm considering submitting the inference of my trained model on test videos but I'm slightly confused here, any help is appreciated! \nQuestion: In order to calculate the log loss we are required to know the ground truth, which we are not provided with for test videos. How do I calculate log loss for test videos? \n\nThanks in advance",
    "730553": "Actually the test set is generated from parts of the full training dataset and you could potentially pull out the ground truth.\n\nOr instead you can separate a portion of your training set (ie 10%) and not train with it and then calculate log loss on that after you train your model.",
    "730566": "Gotcha. That seems doable, but wouldn't that mean all participants have different ways of calculating log loss? On different parts of the data, I mean?",
    "730597": "You need to build a model to predict the probabilities of each sample being a deep fake, this will be a value between 0 and 1, 1.0 being very sure it's a fake, 0 being very sure it's real, 0.5 would mean your model is unsure. Kaggle know the labels and will calculate the log loss of your model's predictions when you submit your solution.",
    "730651": "So you include your model in the inference notebook. When you submit the notebook, the test video dir will be swapped to an unseen validation set. Then, after your model predicts that unseen validation set, it should output as submission.csv. Then kaggle will calculate log loss based on your prediction(submission.csv)",
    "730714": "Perfect! Thank you for the explanations."
  },
  "source": "meta"
}