{
  "id": 132222,
  "title": "Leader Board Scoring Anomaly x3 Identical scores that shouldn't be??",
  "url": "/competitions/deepfake-detection-challenge/discussion/132222",
  "author_name": "",
  "post_date": "2020-02-25T00:22:16.046162600Z",
  "votes": null,
  "comment_count": 7,
  "views": 0,
  "content": "<p>As we only get 2 submissions a day, with no feedback on the run, I get a bit testy when the score is out in left field, not a bad score, if it was an Actual score, I could learn from it.  I suspect it is some form of a scoring issue I am running into and this is some kind of error message or such.</p>\n\n<p>So entering all 0.5 will get a score of (0.69314), which is not a bad score if you are actually trying rather than base lining.  </p>\n\n<p>So I ran a model that limited at 0.8 and 0.2 for a model that both local and commit gave a normal spread of results.  LB Score was (0.69314).</p>\n\n<p>I then tweaked the model with a limit of .75 and 0.25 limits, and got very different, but evenly spread values local and commit.  (of note, I am creating full model with submission, so if training folder was not present it might get all zeros).  again I got (0.69314).</p>\n\n<p>Both of the above due to how the code worked, the results were -1 to 1, that then got scaled to 0-1, so a 0 would have been changed to a 0.5 and I would not be surprised.  </p>\n\n<p>..... time goes by</p>\n\n<p>Just now (a day later), I wasted a submission to take one value that was calculated from the first TEST video that both local and commit had around 0.8 and made the entire label that same value.  (Just a gut check to see what the score would be).  Again I got (0.69314).</p>\n\n<p>Now if the value was 0, I would have gotten 17, if it was anything but 0.500, I would have gotten anything but  (0.69314).  But again I got (0.69314).</p>\n\n<p>So I am mystified, either I have a scoring miracle or there is some LB scoring issue I am running into that gives me a (0.69314) score all the time??</p>\n\n<p>Any ideas?  Is anyone else seeing this?</p>",
  "messages": [
    {
      "id": "755616",
      "postDate": "02/25/2020 00:22:16",
      "content": "<p>As we only get 2 submissions a day, with no feedback on the run, I get a bit testy when the score is out in left field, not a bad score, if it was an Actual score, I could learn from it.  I suspect it is some form of a scoring issue I am running into and this is some kind of error message or such.</p>\n\n<p>So entering all 0.5 will get a score of (0.69314), which is not a bad score if you are actually trying rather than base lining.  </p>\n\n<p>So I ran a model that limited at 0.8 and 0.2 for a model that both local and commit gave a normal spread of results.  LB Score was (0.69314).</p>\n\n<p>I then tweaked the model with a limit of .75 and 0.25 limits, and got very different, but evenly spread values local and commit.  (of note, I am creating full model with submission, so if training folder was not present it might get all zeros).  again I got (0.69314).</p>\n\n<p>Both of the above due to how the code worked, the results were -1 to 1, that then got scaled to 0-1, so a 0 would have been changed to a 0.5 and I would not be surprised.  </p>\n\n<p>..... time goes by</p>\n\n<p>Just now (a day later), I wasted a submission to take one value that was calculated from the first TEST video that both local and commit had around 0.8 and made the entire label that same value.  (Just a gut check to see what the score would be).  Again I got (0.69314).</p>\n\n<p>Now if the value was 0, I would have gotten 17, if it was anything but 0.500, I would have gotten anything but  (0.69314).  But again I got (0.69314).</p>\n\n<p>So I am mystified, either I have a scoring miracle or there is some LB scoring issue I am running into that gives me a (0.69314) score all the time??</p>\n\n<p>Any ideas?  Is anyone else seeing this?</p>",
      "rawMarkdown": "As we only get 2 submissions a day, with no feedback on the run, I get a bit testy when the score is out in left field, not a bad score, if it was an Actual score, I could learn from it.  I suspect it is some form of a scoring issue I am running into and this is some kind of error message or such.\n\nSo entering all 0.5 will get a score of (0.69314), which is not a bad score if you are actually trying rather than base lining.  \n\nSo I ran a model that limited at 0.8 and 0.2 for a model that both local and commit gave a normal spread of results.  LB Score was (0.69314).\n\nI then tweaked the model with a limit of .75 and 0.25 limits, and got very different, but evenly spread values local and commit.  (of note, I am creating full model with submission, so if training folder was not present it might get all zeros).  again I got (0.69314).\n\nBoth of the above due to how the code worked, the results were -1 to 1, that then got scaled to 0-1, so a 0 would have been changed to a 0.5 and I would not be surprised.  \n\n..... time goes by\n\nJust now (a day later), I wasted a submission to take one value that was calculated from the first TEST video that both local and commit had around 0.8 and made the entire label that same value.  (Just a gut check to see what the score would be).  Again I got (0.69314).\n\n\nNow if the value was 0, I would have gotten 17, if it was anything but 0.500, I would have gotten anything but  (0.69314).  But again I got (0.69314).\n\nSo I am mystified, either I have a scoring miracle or there is some LB scoring issue I am running into that gives me a (0.69314) score all the time??\n\nAny ideas?  Is anyone else seeing this?",
      "votes": null
    },
    {
      "id": "755655",
      "postDate": "02/25/2020 01:42:24",
      "content": "<p>You have a bug in your submission pipeline for sure. Your code may read a specific cell, fail, then the default submission file is used for scoring (if you already load submission.csv). If this is not the case, there can be a similar error like that.</p>",
      "rawMarkdown": "You have a bug in your submission pipeline for sure. Your code may read a specific cell, fail, then the default submission file is used for scoring (if you already load submission.csv). If this is not the case, there can be a similar error like that.",
      "votes": null
    },
    {
      "id": "755829",
      "postDate": "02/25/2020 06:44:21",
      "content": "<p>There are 27 corrupt videos in the hidden public test set (the one your submission gets tested against) if your code crashes it would probably put 0.5 for the rest ... I would read the frames with something like this: \n<code>\ntry:\n      v_cap = cv2.VideoCapture(video_Path)\n      for j in range(v_len):\n         success, vframe = v_cap.read()\n         if(success):\n            #Add vframe to your batch or whatever pipelines steps you have\nexcept:\n    raise Exception(\"Stopped at \"+video_Path) \n</code></p>",
      "rawMarkdown": "There are 27 corrupt videos in the hidden public test set (the one your submission gets tested against) if your code crashes it would probably put 0.5 for the rest ... I would read the frames with something like this: \n```\ntry:\n      v_cap = cv2.VideoCapture(video_Path)\n      for j in range(v_len):\n         success, vframe = v_cap.read()\n         if(success):\n            #Add vframe to your batch or whatever pipelines steps you have\nexcept:\n    raise Exception(\"Stopped at \"+video_Path) \n```",
      "votes": null
    },
    {
      "id": "755836",
      "postDate": "02/25/2020 06:58:15",
      "content": "<p>Thanks guys, I can certainly add additional try, except blocks to work around corrupt videos.  I am going a less traveled path, so corrupt videos might have a larger impact.  Thanks much!  I assume the corrupt videos are not flagged as real.</p>",
      "rawMarkdown": "Thanks guys, I can certainly add additional try, except blocks to work around corrupt videos.  I am going a less traveled path, so corrupt videos might have a larger impact.  Thanks much!  I assume the corrupt videos are not flagged as real.",
      "votes": null
    },
    {
      "id": "755864",
      "postDate": "02/25/2020 07:50:05",
      "content": "<p>That is a dangerous thought , it seems they are indeed fake .. but there is no guarantee on the private test set .. the corrupt videos might be the real ones 😄 </p>",
      "rawMarkdown": "That is a dangerous thought , it seems they are indeed fake .. but there is no guarantee on the private test set .. the corrupt videos might be the real ones 😄",
      "votes": null
    },
    {
      "id": "759333",
      "postDate": "02/28/2020 22:41:42",
      "content": "<p>I did flag corrupt videos with 0.9 and the boost I got was from 0.63 to 0.62 so I doubt that will make a huge impact, but try it</p>",
      "rawMarkdown": "I did flag corrupt videos with 0.9 and the boost I got was from 0.63 to 0.62 so I doubt that will make a huge impact, but try it",
      "votes": null
    },
    {
      "id": "759962",
      "postDate": "02/29/2020 16:58:41",
      "content": "<p>So you're saying there may still be a mix of Fake/Real in the corrupt videos?</p>",
      "rawMarkdown": "So you're saying there may still be a mix of Fake/Real in the corrupt videos?",
      "votes": null
    },
    {
      "id": "762892",
      "postDate": "03/03/2020 23:08:45",
      "content": "<p>Just to loop back, Yes, I had initialized my array with 0.5 for later filling with actual prediction values, say Cell 5, then ran model and generated predictions Cells 6-10, then Cell 12 for instance I overwrote the prediction array with the actual predictions and saved to submit file.  So what happened was the corrupted files caused Cells 6-10 to fail/ generate unhanded exception that dropped out, allowing Cell 12 to submit the default value array of 0.5.  </p>\n\n<p>The fix was to manually generate (throw) Exceptions in the file handling calls and make sure they were handled cleanly.  Not just caught, because catching may cause a return value to fail to come back in an acceptable range and cause un-anticipated side failures.  So making sure that in a handled exception case, the return value or state of data is still good for continuing, including not putting values in data that might be in a range none of the other data is in.  I had a -99 error value set in a range that was always positive for instance, that I had to handle better.</p>",
      "rawMarkdown": "Just to loop back, Yes, I had initialized my array with 0.5 for later filling with actual prediction values, say Cell 5, then ran model and generated predictions Cells 6-10, then Cell 12 for instance I overwrote the prediction array with the actual predictions and saved to submit file.  So what happened was the corrupted files caused Cells 6-10 to fail/ generate unhanded exception that dropped out, allowing Cell 12 to submit the default value array of 0.5.  \n\nThe fix was to manually generate (throw) Exceptions in the file handling calls and make sure they were handled cleanly.  Not just caught, because catching may cause a return value to fail to come back in an acceptable range and cause un-anticipated side failures.  So making sure that in a handled exception case, the return value or state of data is still good for continuing, including not putting values in data that might be in a range none of the other data is in.  I had a -99 error value set in a range that was always positive for instance, that I had to handle better.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 755655,
      "author_name": "khahuras",
      "author_url": "",
      "post_date": "02/25/2020 01:42:24",
      "content": "<p>You have a bug in your submission pipeline for sure. Your code may read a specific cell, fail, then the default submission file is used for scoring (if you already load submission.csv). If this is not the case, there can be a similar error like that.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 755829,
      "author_name": "ma7moud",
      "author_url": "",
      "post_date": "02/25/2020 06:44:21",
      "content": "<p>There are 27 corrupt videos in the hidden public test set (the one your submission gets tested against) if your code crashes it would probably put 0.5 for the rest ... I would read the frames with something like this: \n<code>\ntry:\n      v_cap = cv2.VideoCapture(video_Path)\n      for j in range(v_len):\n         success, vframe = v_cap.read()\n         if(success):\n            #Add vframe to your batch or whatever pipelines steps you have\nexcept:\n    raise Exception(\"Stopped at \"+video_Path) \n</code></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 755836,
      "author_name": "meckdahl",
      "author_url": "",
      "post_date": "02/25/2020 06:58:15",
      "content": "<p>Thanks guys, I can certainly add additional try, except blocks to work around corrupt videos.  I am going a less traveled path, so corrupt videos might have a larger impact.  Thanks much!  I assume the corrupt videos are not flagged as real.</p>",
      "votes": null,
      "replies": [
        {
          "id": 755864,
          "author_name": "ma7moud",
          "author_url": "",
          "post_date": "02/25/2020 07:50:05",
          "content": "<p>That is a dangerous thought , it seems they are indeed fake .. but there is no guarantee on the private test set .. the corrupt videos might be the real ones 😄 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 759333,
          "author_name": "basharallabadi",
          "author_url": "",
          "post_date": "02/28/2020 22:41:42",
          "content": "<p>I did flag corrupt videos with 0.9 and the boost I got was from 0.63 to 0.62 so I doubt that will make a huge impact, but try it</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 759962,
          "author_name": "meckdahl",
          "author_url": "",
          "post_date": "02/29/2020 16:58:41",
          "content": "<p>So you're saying there may still be a mix of Fake/Real in the corrupt videos?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 762892,
      "author_name": "meckdahl",
      "author_url": "",
      "post_date": "03/03/2020 23:08:45",
      "content": "<p>Just to loop back, Yes, I had initialized my array with 0.5 for later filling with actual prediction values, say Cell 5, then ran model and generated predictions Cells 6-10, then Cell 12 for instance I overwrote the prediction array with the actual predictions and saved to submit file.  So what happened was the corrupted files caused Cells 6-10 to fail/ generate unhanded exception that dropped out, allowing Cell 12 to submit the default value array of 0.5.  </p>\n\n<p>The fix was to manually generate (throw) Exceptions in the file handling calls and make sure they were handled cleanly.  Not just caught, because catching may cause a return value to fail to come back in an acceptable range and cause un-anticipated side failures.  So making sure that in a handled exception case, the return value or state of data is still good for continuing, including not putting values in data that might be in a range none of the other data is in.  I had a -99 error value set in a range that was always positive for instance, that I had to handle better.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "755616": "As we only get 2 submissions a day, with no feedback on the run, I get a bit testy when the score is out in left field, not a bad score, if it was an Actual score, I could learn from it.  I suspect it is some form of a scoring issue I am running into and this is some kind of error message or such.\n\nSo entering all 0.5 will get a score of (0.69314), which is not a bad score if you are actually trying rather than base lining.  \n\nSo I ran a model that limited at 0.8 and 0.2 for a model that both local and commit gave a normal spread of results.  LB Score was (0.69314).\n\nI then tweaked the model with a limit of .75 and 0.25 limits, and got very different, but evenly spread values local and commit.  (of note, I am creating full model with submission, so if training folder was not present it might get all zeros).  again I got (0.69314).\n\nBoth of the above due to how the code worked, the results were -1 to 1, that then got scaled to 0-1, so a 0 would have been changed to a 0.5 and I would not be surprised.  \n\n..... time goes by\n\nJust now (a day later), I wasted a submission to take one value that was calculated from the first TEST video that both local and commit had around 0.8 and made the entire label that same value.  (Just a gut check to see what the score would be).  Again I got (0.69314).\n\n\nNow if the value was 0, I would have gotten 17, if it was anything but 0.500, I would have gotten anything but  (0.69314).  But again I got (0.69314).\n\nSo I am mystified, either I have a scoring miracle or there is some LB scoring issue I am running into that gives me a (0.69314) score all the time??\n\nAny ideas?  Is anyone else seeing this?",
    "755655": "You have a bug in your submission pipeline for sure. Your code may read a specific cell, fail, then the default submission file is used for scoring (if you already load submission.csv). If this is not the case, there can be a similar error like that.",
    "755829": "There are 27 corrupt videos in the hidden public test set (the one your submission gets tested against) if your code crashes it would probably put 0.5 for the rest ... I would read the frames with something like this: \n```\ntry:\n      v_cap = cv2.VideoCapture(video_Path)\n      for j in range(v_len):\n         success, vframe = v_cap.read()\n         if(success):\n            #Add vframe to your batch or whatever pipelines steps you have\nexcept:\n    raise Exception(\"Stopped at \"+video_Path) \n```",
    "755836": "Thanks guys, I can certainly add additional try, except blocks to work around corrupt videos.  I am going a less traveled path, so corrupt videos might have a larger impact.  Thanks much!  I assume the corrupt videos are not flagged as real.",
    "755864": "That is a dangerous thought , it seems they are indeed fake .. but there is no guarantee on the private test set .. the corrupt videos might be the real ones 😄",
    "759333": "I did flag corrupt videos with 0.9 and the boost I got was from 0.63 to 0.62 so I doubt that will make a huge impact, but try it",
    "759962": "So you're saying there may still be a mix of Fake/Real in the corrupt videos?",
    "762892": "Just to loop back, Yes, I had initialized my array with 0.5 for later filling with actual prediction values, say Cell 5, then ran model and generated predictions Cells 6-10, then Cell 12 for instance I overwrote the prediction array with the actual predictions and saved to submit file.  So what happened was the corrupted files caused Cells 6-10 to fail/ generate unhanded exception that dropped out, allowing Cell 12 to submit the default value array of 0.5.  \n\nThe fix was to manually generate (throw) Exceptions in the file handling calls and make sure they were handled cleanly.  Not just caught, because catching may cause a return value to fail to come back in an acceptable range and cause un-anticipated side failures.  So making sure that in a handled exception case, the return value or state of data is still good for continuing, including not putting values in data that might be in a range none of the other data is in.  I had a -99 error value set in a range that was always positive for instance, that I had to handle better."
  },
  "source": "meta"
}