{
  "id": 122041,
  "title": "[SOLVED] How many test sets do we have?",
  "url": "/competitions/tensorflow2-question-answering/discussion/122041",
  "author_name": "Dmytro Danevskyi",
  "post_date": "2019-12-17T09:36:28.582000",
  "votes": 10,
  "comment_count": 6,
  "views": 0,
  "content": "<p><strong>EDIT</strong>: Everything is fine, just keep in mind that your kernel runs only once on both public + private datasets and then the score is computed only for the public part.</p>\n\n<hr>\n\n<p>I came through a strange issue recently. My current pipeline is not to have full inference code in the Kaggle kernel, but rather to train &amp; predict locally and load the submission file through a private dataset. </p>\n\n<p>In order to handle the private dataset case, I check how many lines are in a <code>sample_submission.csv</code> file, and if the number matches the number found in my submission file, I submit my submission, otherwise I submit the sample submission. Theoretically, this should give approximately zero score for the private dataset, and some non-zero (assuming that my model does its job) score for the public dataset.</p>\n\n<p>However, my kernel gives me EXACTLY 0 score. The resulting submission I see after doing a commit looks exactly as a submission I've generated locally. My submission has a quite significant overlap with other public submission files. But the score is ZERO. </p>\n\n<p>The only explanation I see here is that we actually have an additional public test set that is hidden from us, but in fact, used internally to compute the leaderboard score, so the real public test set (346 examples) is being simply ignored.</p>\n\n<p>I'd appreciate it if anyone could shed some light on this problem. <a href=\"/philculliton\">@philculliton</a> could you provide some feedback on this?</p>\n\n<p>My problem seems to overlap with <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/119280\">this topic</a>, but the topic doesn't seem to provide an answer.</p>",
  "messages": [
    {
      "id": 696954,
      "postDate": "2019-12-17T09:36:28.583Z",
      "content": "<p><strong>EDIT</strong>: Everything is fine, just keep in mind that your kernel runs only once on both public + private datasets and then the score is computed only for the public part.</p>\n\n<hr>\n\n<p>I came through a strange issue recently. My current pipeline is not to have full inference code in the Kaggle kernel, but rather to train &amp; predict locally and load the submission file through a private dataset. </p>\n\n<p>In order to handle the private dataset case, I check how many lines are in a <code>sample_submission.csv</code> file, and if the number matches the number found in my submission file, I submit my submission, otherwise I submit the sample submission. Theoretically, this should give approximately zero score for the private dataset, and some non-zero (assuming that my model does its job) score for the public dataset.</p>\n\n<p>However, my kernel gives me EXACTLY 0 score. The resulting submission I see after doing a commit looks exactly as a submission I've generated locally. My submission has a quite significant overlap with other public submission files. But the score is ZERO. </p>\n\n<p>The only explanation I see here is that we actually have an additional public test set that is hidden from us, but in fact, used internally to compute the leaderboard score, so the real public test set (346 examples) is being simply ignored.</p>\n\n<p>I'd appreciate it if anyone could shed some light on this problem. <a href=\"/philculliton\">@philculliton</a> could you provide some feedback on this?</p>\n\n<p>My problem seems to overlap with <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/119280\">this topic</a>, but the topic doesn't seem to provide an answer.</p>",
      "rawMarkdown": "**EDIT**: Everything is fine, just keep in mind that your kernel runs only once on both public + private datasets and then the score is computed only for the public part.\n\n\n------\n\n\nI came through a strange issue recently. My current pipeline is not to have full inference code in the Kaggle kernel, but rather to train &amp; predict locally and load the submission file through a private dataset. \n\nIn order to handle the private dataset case, I check how many lines are in a `sample_submission.csv` file, and if the number matches the number found in my submission file, I submit my submission, otherwise I submit the sample submission. Theoretically, this should give approximately zero score for the private dataset, and some non-zero (assuming that my model does its job) score for the public dataset.\n\nHowever, my kernel gives me EXACTLY 0 score. The resulting submission I see after doing a commit looks exactly as a submission I've generated locally. My submission has a quite significant overlap with other public submission files. But the score is ZERO. \n\nThe only explanation I see here is that we actually have an additional public test set that is hidden from us, but in fact, used internally to compute the leaderboard score, so the real public test set (346 examples) is being simply ignored.\n\nI'd appreciate it if anyone could shed some light on this problem. @philculliton could you provide some feedback on this?\n\nMy problem seems to overlap with [this topic](https://www.kaggle.com/c/tensorflow2-question-answering/discussion/119280), but the topic doesn't seem to provide an answer.\n",
      "votes": 10
    },
    {
      "id": 697026,
      "postDate": "2019-12-17T11:39:55.303Z",
      "content": "<p>Indeed, it's not only the metric which is obscure but the submission process as well. </p>\n\n<p>As Phil helped me a bit with debugging, I'll clarify on the issue that I encountered (though it's mentioned in <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/119280\">the same post</a>).</p>\n\n<p>I used a <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/118129\">similar hack</a> to save time (not working for me anymore, <code>Notebook Threw Exception</code>). The problem with my code was that it worked fine with public test data (346-long) but failed on the private part (which is from 3k to 3500-long), and I actually figure out why. </p>\n\n<p>Turned out that after an exception thrown for the private test part, the submission \"backed off\" to sample submission, therefore Zeros. When I executed the buggy code without hacks saving time, I got <code>Submission CSV Not Found</code>. </p>",
      "rawMarkdown": "Indeed, it's not only the metric which is obscure but the submission process as well. \n\nAs Phil helped me a bit with debugging, I'll clarify on the issue that I encountered (though it's mentioned in [the same post](https://www.kaggle.com/c/tensorflow2-question-answering/discussion/119280)).\n\nI used a [similar hack](https://www.kaggle.com/c/tensorflow2-question-answering/discussion/118129) to save time (not working for me anymore, `Notebook Threw Exception`). The problem with my code was that it worked fine with public test data (346-long) but failed on the private part (which is from 3k to 3500-long), and I actually figure out why. \n\nTurned out that after an exception thrown for the private test part, the submission \"backed off\" to sample submission, therefore Zeros. When I executed the buggy code without hacks saving time, I got `Submission CSV Not Found`. \n\n\n",
      "votes": 2,
      "replies": [
        {
          "id": 697095,
          "postDate": "2019-12-17T13:28:48.650Z",
          "content": "<p>Today, I also got <code>Submission CSV Not Found</code>. And when I just sent all blank answer (without any prediction from model), I got <code>Notebook Threw Exception</code>. Commit is working, but not submission.</p>",
          "rawMarkdown": "Today, I also got `Submission CSV Not Found`. And when I just sent all blank answer (without any prediction from model), I got `Notebook Threw Exception`. Commit is working, but not submission."
        }
      ]
    },
    {
      "id": 697024,
      "postDate": "2019-12-17T11:34:31.333Z",
      "content": "<p>Are you doing like this? If so, your submission always use sample submission.</p>\n\n<p><code>\nif len(submission_df) == 346 * 2:\n    # submit your submission.csv\nelse:\n    # submit sample_submission.csv\n</code></p>\n\n<p>I tried same process and found that the <code>if len(submission_df) == 346 * 2</code> statement is satified only in COMMIT and not in SUBMISSION.\nI guess the submission process include both public and private test set, then compute public LB by using public test set only. </p>",
      "rawMarkdown": "Are you doing like this? If so, your submission always use sample submission.\n\n```\nif len(submission_df) == 346 * 2:\n    # submit your submission.csv\nelse:\n    # submit sample_submission.csv\n```\n\nI tried same process and found that the `if len(submission_df) == 346 * 2` statement is satified only in COMMIT and not in SUBMISSION.\nI guess the submission process include both public and private test set, then compute public LB by using public test set only. ",
      "votes": 2,
      "replies": [
        {
          "id": 697037,
          "postDate": "2019-12-17T12:04:25.067Z",
          "content": "<p>You are right. The proper way of doing this is to fill the sample submission file with the predictions made locally, not to replace the file.</p>",
          "rawMarkdown": "You are right. The proper way of doing this is to fill the sample submission file with the predictions made locally, not to replace the file.",
          "votes": 1
        },
        {
          "id": 697055,
          "postDate": "2019-12-17T12:23:10.433Z",
          "content": "<p>Yes, at SUBMISSION stage <code>len(submission_df)</code> is now about 7k-long so you always submit a sample submission file. </p>",
          "rawMarkdown": "Yes, at SUBMISSION stage `len(submission_df)` is now about 7k-long so you always submit a sample submission file. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 697104,
      "postDate": "2019-12-17T13:42:52.510Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 697026,
      "author_name": "Yury Kashnitsky",
      "author_url": "",
      "post_date": "2019-12-17T11:39:55.303000",
      "content": "<p>Indeed, it's not only the metric which is obscure but the submission process as well. </p>\n\n<p>As Phil helped me a bit with debugging, I'll clarify on the issue that I encountered (though it's mentioned in <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/119280\">the same post</a>).</p>\n\n<p>I used a <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/118129\">similar hack</a> to save time (not working for me anymore, <code>Notebook Threw Exception</code>). The problem with my code was that it worked fine with public test data (346-long) but failed on the private part (which is from 3k to 3500-long), and I actually figure out why. </p>\n\n<p>Turned out that after an exception thrown for the private test part, the submission \"backed off\" to sample submission, therefore Zeros. When I executed the buggy code without hacks saving time, I got <code>Submission CSV Not Found</code>. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 697095,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2019-12-17T13:28:48.650000",
          "content": "<p>Today, I also got <code>Submission CSV Not Found</code>. And when I just sent all blank answer (without any prediction from model), I got <code>Notebook Threw Exception</code>. Commit is working, but not submission.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 697024,
      "author_name": "cfiken",
      "author_url": "",
      "post_date": "2019-12-17T11:34:31.333000",
      "content": "<p>Are you doing like this? If so, your submission always use sample submission.</p>\n\n<p><code>\nif len(submission_df) == 346 * 2:\n    # submit your submission.csv\nelse:\n    # submit sample_submission.csv\n</code></p>\n\n<p>I tried same process and found that the <code>if len(submission_df) == 346 * 2</code> statement is satified only in COMMIT and not in SUBMISSION.\nI guess the submission process include both public and private test set, then compute public LB by using public test set only. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 697037,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2019-12-17T12:04:25.067000",
          "content": "<p>You are right. The proper way of doing this is to fill the sample submission file with the predictions made locally, not to replace the file.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 697055,
          "author_name": "Yury Kashnitsky",
          "author_url": "",
          "post_date": "2019-12-17T12:23:10.433000",
          "content": "<p>Yes, at SUBMISSION stage <code>len(submission_df)</code> is now about 7k-long so you always submit a sample submission file. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 697104,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-17T13:42:52.510000",
      "content": "",
      "votes": -1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "696954": "**EDIT**: Everything is fine, just keep in mind that your kernel runs only once on both public + private datasets and then the score is computed only for the public part.\n\n\n------\n\n\nI came through a strange issue recently. My current pipeline is not to have full inference code in the Kaggle kernel, but rather to train &amp; predict locally and load the submission file through a private dataset. \n\nIn order to handle the private dataset case, I check how many lines are in a `sample_submission.csv` file, and if the number matches the number found in my submission file, I submit my submission, otherwise I submit the sample submission. Theoretically, this should give approximately zero score for the private dataset, and some non-zero (assuming that my model does its job) score for the public dataset.\n\nHowever, my kernel gives me EXACTLY 0 score. The resulting submission I see after doing a commit looks exactly as a submission I've generated locally. My submission has a quite significant overlap with other public submission files. But the score is ZERO. \n\nThe only explanation I see here is that we actually have an additional public test set that is hidden from us, but in fact, used internally to compute the leaderboard score, so the real public test set (346 examples) is being simply ignored.\n\nI'd appreciate it if anyone could shed some light on this problem. @philculliton could you provide some feedback on this?\n\nMy problem seems to overlap with [this topic](https://www.kaggle.com/c/tensorflow2-question-answering/discussion/119280), but the topic doesn't seem to provide an answer.\n",
    "697026": "Indeed, it's not only the metric which is obscure but the submission process as well. \n\nAs Phil helped me a bit with debugging, I'll clarify on the issue that I encountered (though it's mentioned in [the same post](https://www.kaggle.com/c/tensorflow2-question-answering/discussion/119280)).\n\nI used a [similar hack](https://www.kaggle.com/c/tensorflow2-question-answering/discussion/118129) to save time (not working for me anymore, `Notebook Threw Exception`). The problem with my code was that it worked fine with public test data (346-long) but failed on the private part (which is from 3k to 3500-long), and I actually figure out why. \n\nTurned out that after an exception thrown for the private test part, the submission \"backed off\" to sample submission, therefore Zeros. When I executed the buggy code without hacks saving time, I got `Submission CSV Not Found`. \n\n\n",
    "697024": "Are you doing like this? If so, your submission always use sample submission.\n\n```\nif len(submission_df) == 346 * 2:\n    # submit your submission.csv\nelse:\n    # submit sample_submission.csv\n```\n\nI tried same process and found that the `if len(submission_df) == 346 * 2` statement is satified only in COMMIT and not in SUBMISSION.\nI guess the submission process include both public and private test set, then compute public LB by using public test set only. ",
    "697104": ""
  }
}