{
  "id": 145947,
  "title": "Public set size 4k Private Set Size 10k == Notebook timeout",
  "url": "/competitions/deepfake-detection-challenge/discussion/145947",
  "author_name": "",
  "post_date": "2020-04-25T07:17:55.539652900Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>According to organisers:</p>\n\n<blockquote>\n  <p>There was a ~6.5% failure rate on re-runs, due mostly to incomplete submissions (submissions which failed to create predictions for all samples). <strong>We were all frankly very impressed by that number</strong> - thanks to everyone for your hard work and attention to detail.</p>\n</blockquote>\n\n<p>They were <strong>very impressed</strong> because the public set size was 4000 images. The private set size was <strong>10,000</strong>. If you assumed they would be the same size your notebook probably timed out before the 9 hour limit.</p>\n\n<p>By their own admission they were <strong>very surprised</strong> more submissions didn't fail. So they were deliberate in their approach. It seems this was a game for the organisers and somehow ppl knew the private set size would be much larger than the public test set.</p>\n\n<p>According to the rules they drop a hint:</p>\n\n<blockquote>\n  <p>Notebook Timeout: Your submission notebook exceeded the allowed runtime cap for the competition. Review the competition's Code Requirements for details, and note that the privately held rerun test set <strong>could be orders of magnitude larger than the publicly shared validation set</strong>.</p>\n</blockquote>\n\n<p>How ppl accounted for this we will never know. Either way, all it did was disqualify teams that may have produced useful models and now the organisers will never know. To me that's just throwing $'s down the toilet.</p>",
  "messages": [
    {
      "id": "820144",
      "postDate": "04/25/2020 07:17:55",
      "content": "<p>According to organisers:</p>\n\n<blockquote>\n  <p>There was a ~6.5% failure rate on re-runs, due mostly to incomplete submissions (submissions which failed to create predictions for all samples). <strong>We were all frankly very impressed by that number</strong> - thanks to everyone for your hard work and attention to detail.</p>\n</blockquote>\n\n<p>They were <strong>very impressed</strong> because the public set size was 4000 images. The private set size was <strong>10,000</strong>. If you assumed they would be the same size your notebook probably timed out before the 9 hour limit.</p>\n\n<p>By their own admission they were <strong>very surprised</strong> more submissions didn't fail. So they were deliberate in their approach. It seems this was a game for the organisers and somehow ppl knew the private set size would be much larger than the public test set.</p>\n\n<p>According to the rules they drop a hint:</p>\n\n<blockquote>\n  <p>Notebook Timeout: Your submission notebook exceeded the allowed runtime cap for the competition. Review the competition's Code Requirements for details, and note that the privately held rerun test set <strong>could be orders of magnitude larger than the publicly shared validation set</strong>.</p>\n</blockquote>\n\n<p>How ppl accounted for this we will never know. Either way, all it did was disqualify teams that may have produced useful models and now the organisers will never know. To me that's just throwing $'s down the toilet.</p>",
      "rawMarkdown": "According to organisers:\n\n&gt; There was a ~6.5% failure rate on re-runs, due mostly to incomplete submissions (submissions which failed to create predictions for all samples). **We were all frankly very impressed by that number** - thanks to everyone for your hard work and attention to detail.\n\nThey were **very impressed** because the public set size was 4000 images. The private set size was **10,000**. If you assumed they would be the same size your notebook probably timed out before the 9 hour limit.\n\nBy their own admission they were **very surprised** more submissions didn't fail. So they were deliberate in their approach. It seems this was a game for the organisers and somehow ppl knew the private set size would be much larger than the public test set.\n\nAccording to the rules they drop a hint:\n\n&gt; Notebook Timeout: Your submission notebook exceeded the allowed runtime cap for the competition. Review the competition's Code Requirements for details, and note that the privately held rerun test set **could be orders of magnitude larger than the publicly shared validation set**.\n\nHow ppl accounted for this we will never know. Either way, all it did was disqualify teams that may have produced useful models and now the organisers will never know. To me that's just throwing $'s down the toilet.",
      "votes": null
    },
    {
      "id": "820226",
      "postDate": "04/25/2020 08:55:48",
      "content": "<blockquote>\n  <p>If you assumed they would be the same size your notebook probably timed out before the 9 hour limit.</p>\n</blockquote>\n\n<p>Where does it say the private set also has a 9 hour limit? </p>",
      "rawMarkdown": "&gt; If you assumed they would be the same size your notebook probably timed out before the 9 hour limit.\n\nWhere does it say the private set also has a 9 hour limit?",
      "votes": null
    },
    {
      "id": "820238",
      "postDate": "04/25/2020 09:10:33",
      "content": "<p>This is what they posted.</p>\n\n<blockquote>\n  <p>Here are the reasons you may have selected submissions which were invalidated, in rough order of how common the issue was:\n  Submission produced no or null predictions on some portion of the test set\n  Submission produced predictions but with the wrong number of rows\n  Your notebooks or linked notebooks/datasets for selected submissions were deleted, and therefore we could not re-run them\n  Insufficient error handling &amp; Misc code failures (including out of memory &amp; <strong>timeout errors</strong>)</p>\n</blockquote>\n\n<p>if there are timeout errors then there has to be a limit. I presume that limit is the same 9 hours.</p>",
      "rawMarkdown": "This is what they posted.\n\n&gt; Here are the reasons you may have selected submissions which were invalidated, in rough order of how common the issue was:\nSubmission produced no or null predictions on some portion of the test set\nSubmission produced predictions but with the wrong number of rows\nYour notebooks or linked notebooks/datasets for selected submissions were deleted, and therefore we could not re-run them\nInsufficient error handling &amp; Misc code failures (including out of memory &amp; **timeout errors**)\n\nif there are timeout errors then there has to be a limit. I presume that limit is the same 9 hours.",
      "votes": null
    },
    {
      "id": "820279",
      "postDate": "04/25/2020 09:43:39",
      "content": "<p>Our kernels, and I assume everyone else's too, were taking just less than 9 hrs to complete (e.g. 8hrs and 30min). There's no way it would run 10k videos within 9hrs.</p>",
      "rawMarkdown": "Our kernels, and I assume everyone else's too, were taking just less than 9 hrs to complete (e.g. 8hrs and 30min). There's no way it would run 10k videos within 9hrs.",
      "votes": null
    },
    {
      "id": "820364",
      "postDate": "04/25/2020 11:36:54",
      "content": "<blockquote>\n  <p>I presume that limit is the same 9 hours.</p>\n</blockquote>\n\n<p>I don't remember in which thread, but I recall the organizers at some point posted that the limits would be scaled up proportionally with the size of the private test set.</p>",
      "rawMarkdown": "&gt; I presume that limit is the same 9 hours.\n\nI don't remember in which thread, but I recall the organizers at some point posted that the limits would be scaled up proportionally with the size of the private test set.",
      "votes": null
    },
    {
      "id": "820412",
      "postDate": "04/25/2020 12:25:49",
      "content": "<p>Well I’m lost then. Just trying to figure out what went wrong with no help from organisers.</p>",
      "rawMarkdown": "Well I’m lost then. Just trying to figure out what went wrong with no help from organisers.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 820226,
      "author_name": "humananalog",
      "author_url": "",
      "post_date": "04/25/2020 08:55:48",
      "content": "<blockquote>\n  <p>If you assumed they would be the same size your notebook probably timed out before the 9 hour limit.</p>\n</blockquote>\n\n<p>Where does it say the private set also has a 9 hour limit? </p>",
      "votes": null,
      "replies": [
        {
          "id": 820238,
          "author_name": "maralski",
          "author_url": "",
          "post_date": "04/25/2020 09:10:33",
          "content": "<p>This is what they posted.</p>\n\n<blockquote>\n  <p>Here are the reasons you may have selected submissions which were invalidated, in rough order of how common the issue was:\n  Submission produced no or null predictions on some portion of the test set\n  Submission produced predictions but with the wrong number of rows\n  Your notebooks or linked notebooks/datasets for selected submissions were deleted, and therefore we could not re-run them\n  Insufficient error handling &amp; Misc code failures (including out of memory &amp; <strong>timeout errors</strong>)</p>\n</blockquote>\n\n<p>if there are timeout errors then there has to be a limit. I presume that limit is the same 9 hours.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 820279,
          "author_name": "hmendonca",
          "author_url": "",
          "post_date": "04/25/2020 09:43:39",
          "content": "<p>Our kernels, and I assume everyone else's too, were taking just less than 9 hrs to complete (e.g. 8hrs and 30min). There's no way it would run 10k videos within 9hrs.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 820364,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "04/25/2020 11:36:54",
          "content": "<blockquote>\n  <p>I presume that limit is the same 9 hours.</p>\n</blockquote>\n\n<p>I don't remember in which thread, but I recall the organizers at some point posted that the limits would be scaled up proportionally with the size of the private test set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 820412,
          "author_name": "maralski",
          "author_url": "",
          "post_date": "04/25/2020 12:25:49",
          "content": "<p>Well I’m lost then. Just trying to figure out what went wrong with no help from organisers.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "820144": "According to organisers:\n\n&gt; There was a ~6.5% failure rate on re-runs, due mostly to incomplete submissions (submissions which failed to create predictions for all samples). **We were all frankly very impressed by that number** - thanks to everyone for your hard work and attention to detail.\n\nThey were **very impressed** because the public set size was 4000 images. The private set size was **10,000**. If you assumed they would be the same size your notebook probably timed out before the 9 hour limit.\n\nBy their own admission they were **very surprised** more submissions didn't fail. So they were deliberate in their approach. It seems this was a game for the organisers and somehow ppl knew the private set size would be much larger than the public test set.\n\nAccording to the rules they drop a hint:\n\n&gt; Notebook Timeout: Your submission notebook exceeded the allowed runtime cap for the competition. Review the competition's Code Requirements for details, and note that the privately held rerun test set **could be orders of magnitude larger than the publicly shared validation set**.\n\nHow ppl accounted for this we will never know. Either way, all it did was disqualify teams that may have produced useful models and now the organisers will never know. To me that's just throwing $'s down the toilet.",
    "820226": "&gt; If you assumed they would be the same size your notebook probably timed out before the 9 hour limit.\n\nWhere does it say the private set also has a 9 hour limit?",
    "820238": "This is what they posted.\n\n&gt; Here are the reasons you may have selected submissions which were invalidated, in rough order of how common the issue was:\nSubmission produced no or null predictions on some portion of the test set\nSubmission produced predictions but with the wrong number of rows\nYour notebooks or linked notebooks/datasets for selected submissions were deleted, and therefore we could not re-run them\nInsufficient error handling &amp; Misc code failures (including out of memory &amp; **timeout errors**)\n\nif there are timeout errors then there has to be a limit. I presume that limit is the same 9 hours.",
    "820279": "Our kernels, and I assume everyone else's too, were taking just less than 9 hrs to complete (e.g. 8hrs and 30min). There's no way it would run 10k videos within 9hrs.",
    "820364": "&gt; I presume that limit is the same 9 hours.\n\nI don't remember in which thread, but I recall the organizers at some point posted that the limits would be scaled up proportionally with the size of the private test set.",
    "820412": "Well I’m lost then. Just trying to figure out what went wrong with no help from organisers."
  },
  "source": "meta"
}