{
  "id": 139518,
  "title": "Getting a submission scoring error",
  "url": "/competitions/deepfake-detection-challenge/discussion/139518",
  "author_name": "Debopriyo Sanyal",
  "post_date": "2020-03-29T06:27:06.260000",
  "votes": 0,
  "comment_count": 23,
  "views": 0,
  "content": "<p>I am getting a submission scoring error, the code works okay but whenever I submit it gives me a scoring error.</p>",
  "messages": [
    {
      "id": 794449,
      "postDate": "2020-04-01T20:34:11.630Z",
      "content": "<p>I've got submission scoring errors on notebooks that ran fine before, maybe that was because their run on public lb has finished past the deadline (though I've submitted those a few hours before it) Will have to wait for late submission to open to check it.</p>",
      "rawMarkdown": "I've got submission scoring errors on notebooks that ran fine before, maybe that was because their run on public lb has finished past the deadline (though I've submitted those a few hours before it) Will have to wait for late submission to open to check it.",
      "replies": [
        {
          "id": 794833,
          "postDate": "2020-04-02T06:01:43.257Z",
          "content": "<p>@petya Are you saying that now late submission would be enabled, past the deadline?</p>",
          "rawMarkdown": "@petya Are you saying that now late submission would be enabled, past the deadline?"
        },
        {
          "id": 794837,
          "postDate": "2020-04-02T06:05:40.090Z",
          "content": "<p><a href=\"/ds123456\">@ds123456</a> well, you can usually submit your results after the competition's private leaderboard has been published. You are not going to be displayed in leaderboards but you can still check how your submission would've performed in the public leaderboard.</p>",
          "rawMarkdown": "@ds123456 well, you can usually submit your results after the competition's private leaderboard has been published. You are not going to be displayed in leaderboards but you can still check how your submission would've performed in the public leaderboard."
        }
      ]
    },
    {
      "id": 793518,
      "postDate": "2020-04-01T04:41:56.627Z",
      "content": "<p>Can someone tell what went wrong here:</p>\n\n<p><a href=\"https://www.kaggle.com/ds123456/nuclear-zone?scriptVersionId=31176407\">https://www.kaggle.com/ds123456/nuclear-zone?scriptVersionId=31176407</a></p>",
      "rawMarkdown": "Can someone tell what went wrong here:\n\nhttps://www.kaggle.com/ds123456/nuclear-zone?scriptVersionId=31176407",
      "replies": [
        {
          "id": 793606,
          "postDate": "2020-04-01T06:30:42.973Z",
          "content": "<p>hi dude, i looked into yours. 2 points</p>\n\n<p><strong>1) memory overflow?</strong></p>\n\n<p>noticed that your way of generating 'frame1' consumed RAM upto 3.5GB with 400 files. this amount would be accumulated by the number of mp4 videos loaded, i'd say that it would exceed 4.9GB limit when you use it for private test set (4000 vids). so I, as other top rankers also disclosed, would like you to use ffmpeg.</p>\n\n<p><strong>2) generating submission.csv</strong></p>\n\n<p>since my way of writing codes are largely different from yours, let me provide simpler way in generating submission.csv as an example. try to replace your with this kind below.</p>\n\n<pre><code># predict with pretrained tf model first, and store predictions in a dict named 'result'\n    for mp4_filename in mp4_filename_list:\n        preds = model.predict_proba(x)\n        result[mp4_filename] = np.squeeze(preds[0][0])\n\n# generate submission.csv next.\nsample_submission = pd.read_csv('../input/deepfake-detection-challenge/sample_submission.csv')\nfor c in sample_submission.columns[sample_submission.columns == 'label']:\n    try:\n        sample_submission[c] = result.values()    # writing corresponding predicted value here\n    except:\n        sample_submission[c] = 0.5    # if no predicted value found, allocate 0.5 \nsample_submission.to_csv('submission.csv', index=False)\n</code></pre>\n\n<p>Nice if these above works for you even a bit. Let's see in the next competition again. :-)</p>",
          "rawMarkdown": "hi dude, i looked into yours. 2 points\n\n**1) memory overflow?**\n\nnoticed that your way of generating 'frame1' consumed RAM upto 3.5GB with 400 files. this amount would be accumulated by the number of mp4 videos loaded, i'd say that it would exceed 4.9GB limit when you use it for private test set (4000 vids). so I, as other top rankers also disclosed, would like you to use ffmpeg.\n\n**2) generating submission.csv**\n\nsince my way of writing codes are largely different from yours, let me provide simpler way in generating submission.csv as an example. try to replace your with this kind below.\n\n    # predict with pretrained tf model first, and store predictions in a dict named 'result'\n        for mp4_filename in mp4_filename_list:\n            preds = model.predict_proba(x)\n            result[mp4_filename] = np.squeeze(preds[0][0])\n\n    # generate submission.csv next.\n    sample_submission = pd.read_csv('../input/deepfake-detection-challenge/sample_submission.csv')\n    for c in sample_submission.columns[sample_submission.columns == 'label']:\n        try:\n            sample_submission[c] = result.values()    # writing corresponding predicted value here\n        except:\n            sample_submission[c] = 0.5    # if no predicted value found, allocate 0.5 \n    sample_submission.to_csv('submission.csv', index=False)\n\nNice if these above works for you even a bit. Let's see in the next competition again. :-)",
          "votes": 1
        },
        {
          "id": 793819,
          "postDate": "2020-04-01T10:11:32.490Z",
          "content": "<p>You should read carefully the competition getting started tab. Your code has 400 hard coded. When you submit, the test videos dir is replaced with a new one that has 4000 entries </p>",
          "rawMarkdown": "You should read carefully the competition getting started tab. Your code has 400 hard coded. When you submit, the test videos dir is replaced with a new one that has 4000 entries ",
          "votes": 1
        },
        {
          "id": 793820,
          "postDate": "2020-04-01T10:12:14.810Z",
          "content": "<p>Also, you should add proper error handling. Some of the 4000 videos are corrupted </p>",
          "rawMarkdown": "Also, you should add proper error handling. Some of the 4000 videos are corrupted ",
          "votes": 1
        },
        {
          "id": 794029,
          "postDate": "2020-04-01T13:53:34.890Z",
          "content": "<p>@Moshel Thanks for your observations.After you mentioned I checked the getting started tab again: 1. I did not see a mention of 4000 entries anywhere. 2. Regardless; I agree the coding should have been so done so as to handle any number of files and not just 400.3. I should have put in a condition to handle corrupted files, this just did not occur to me.</p>",
          "rawMarkdown": "@Moshel Thanks for your observations.After you mentioned I checked the getting started tab again: 1. I did not see a mention of 4000 entries anywhere. 2. Regardless; I agree the coding should have been so done so as to handle any number of files and not just 400.3. I should have put in a condition to handle corrupted files, this just did not occur to me."
        }
      ]
    },
    {
      "id": 790981,
      "postDate": "2020-03-30T02:26:54.003Z",
      "content": "<p>I am getting a submission error as well - does this take away from my available submissions?</p>",
      "rawMarkdown": "I am getting a submission error as well - does this take away from my available submissions?"
    },
    {
      "id": 790500,
      "postDate": "2020-03-29T16:24:55.790Z",
      "content": "<p>Make sure you do not save intermediate files to disk, kernel has 5GB of storage for output/working dir.</p>",
      "rawMarkdown": "Make sure you do not save intermediate files to disk, kernel has 5GB of storage for output/working dir."
    },
    {
      "id": 790086,
      "postDate": "2020-03-29T09:24:28.723Z",
      "content": "<p>I had this error in different 2 cases:\n- I forgot the csv header, I wrote only the results.\n- I had nan value at some results. Check zero division, etc.</p>",
      "rawMarkdown": "I had this error in different 2 cases:\n- I forgot the csv header, I wrote only the results.\n- I had nan value at some results. Check zero division, etc.",
      "replies": [
        {
          "id": 790138,
          "postDate": "2020-03-29T10:29:36.533Z",
          "content": "<p>I am using  df.to_csv(\"submission.csv\", encoding='utf-8', index=False) to print my output to file. Would doing that create a csv header in the generated file? My understanding is that it does, however you know correct me if I am ...</p>",
          "rawMarkdown": "I am using  df.to_csv(\"submission.csv\", encoding='utf-8', index=False) to print my output to file. Would doing that create a csv header in the generated file? My understanding is that it does, however you know correct me if I am ..."
        }
      ]
    },
    {
      "id": 790032,
      "postDate": "2020-03-29T08:09:29.377Z",
      "content": "<p>ds, I'm also suffering from same challenge (another topic was made just now). How much time does your notebook spend to run with 400 public test videos? I'm asking this because I heard that roughly 10 times of the time spent for public test data needs to be below 9 hours. Mine does 40 mins. If anything rings a bell with you, pls try check it out.   Mats</p>",
      "rawMarkdown": "ds, I'm also suffering from same challenge (another topic was made just now). How much time does your notebook spend to run with 400 public test videos? I'm asking this because I heard that roughly 10 times of the time spent for public test data needs to be below 9 hours. Mine does 40 mins. If anything rings a bell with you, pls try check it out.   Mats",
      "replies": [
        {
          "id": 790050,
          "postDate": "2020-03-29T08:31:55.487Z",
          "content": "<p>my notebook when run cell by cell takes less than 30 mins.</p>",
          "rawMarkdown": "my notebook when run cell by cell takes less than 30 mins."
        },
        {
          "id": 790058,
          "postDate": "2020-03-29T08:41:18.437Z",
          "content": "<p>ok. got it. just minutes ago, one smart person suggested me to output NOT '1, 0, 0, 1, ...' kind of value to submission.csv but to do actual float value. Mie was done only with integers. If your submission.csv also contains such integer only, worth trying it if you would like to. I'll keep following this topic.  I'm running my code now. Mats</p>",
          "rawMarkdown": "ok. got it. just minutes ago, one smart person suggested me to output NOT '1, 0, 0, 1, ...' kind of value to submission.csv but to do actual float value. Mie was done only with integers. If your submission.csv also contains such integer only, worth trying it if you would like to. I'll keep following this topic.  I'm running my code now. Mats"
        },
        {
          "id": 790124,
          "postDate": "2020-03-29T10:08:45.420Z",
          "content": "<p>Did the float thing help solve your issue?</p>",
          "rawMarkdown": "Did the float thing help solve your issue?"
        },
        {
          "id": 790132,
          "postDate": "2020-03-29T10:22:33.980Z",
          "content": "<p>still working on it. will keep u posted when done.</p>",
          "rawMarkdown": "still working on it. will keep u posted when done."
        },
        {
          "id": 790972,
          "postDate": "2020-03-30T02:02:28.207Z",
          "content": "<p>hey buddy, float didn't work.</p>\n\n<p><a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/139528\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/139528</a></p>",
          "rawMarkdown": "hey buddy, float didn't work.\n\nhttps://www.kaggle.com/c/deepfake-detection-challenge/discussion/139528"
        }
      ]
    },
    {
      "id": 789946,
      "postDate": "2020-03-29T06:30:37.007Z",
      "content": "<p>Any ideas how to go about this submission scoring error? The code works okay when run in notebook.</p>",
      "rawMarkdown": "Any ideas how to go about this submission scoring error? The code works okay when run in notebook.",
      "replies": [
        {
          "id": 789983,
          "postDate": "2020-03-29T07:21:34.920Z",
          "content": "<p>Could you specify what the error says or a screenshot?</p>",
          "rawMarkdown": "Could you specify what the error says or a screenshot?"
        },
        {
          "id": 790023,
          "postDate": "2020-03-29T08:03:17.857Z",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3369259%2F1346a7aa524460851f09a694865653d2%2FScreenshot%202020-03-29%20at%201.32.19%20PM.png?generation=1585468990178305&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3369259%2F1346a7aa524460851f09a694865653d2%2FScreenshot%202020-03-29%20at%201.32.19%20PM.png?generation=1585468990178305&amp;alt=media)\n"
        },
        {
          "id": 790024,
          "postDate": "2020-03-29T08:03:44.760Z",
          "content": "<p>Please find above the error screenshot</p>",
          "rawMarkdown": "Please find above the error screenshot"
        }
      ]
    },
    {
      "id": 789942,
      "postDate": "2020-03-29T06:27:06.260Z",
      "content": "<p>I am getting a submission scoring error, the code works okay but whenever I submit it gives me a scoring error.</p>",
      "rawMarkdown": "I am getting a submission scoring error, the code works okay but whenever I submit it gives me a scoring error."
    },
    {
      "id": 792597,
      "postDate": "2020-03-31T11:04:46.350Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 794449,
      "author_name": "petya",
      "author_url": "",
      "post_date": "2020-04-01T20:34:11.630000",
      "content": "<p>I've got submission scoring errors on notebooks that ran fine before, maybe that was because their run on public lb has finished past the deadline (though I've submitted those a few hours before it) Will have to wait for late submission to open to check it.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 794833,
          "author_name": "Debopriyo Sanyal",
          "author_url": "",
          "post_date": "2020-04-02T06:01:43.257000",
          "content": "<p>@petya Are you saying that now late submission would be enabled, past the deadline?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 794837,
          "author_name": "petya",
          "author_url": "",
          "post_date": "2020-04-02T06:05:40.090000",
          "content": "<p><a href=\"/ds123456\">@ds123456</a> well, you can usually submit your results after the competition's private leaderboard has been published. You are not going to be displayed in leaderboards but you can still check how your submission would've performed in the public leaderboard.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 793518,
      "author_name": "Debopriyo Sanyal",
      "author_url": "",
      "post_date": "2020-04-01T04:41:56.627000",
      "content": "<p>Can someone tell what went wrong here:</p>\n\n<p><a href=\"https://www.kaggle.com/ds123456/nuclear-zone?scriptVersionId=31176407\">https://www.kaggle.com/ds123456/nuclear-zone?scriptVersionId=31176407</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 793606,
          "author_name": "Mats",
          "author_url": "",
          "post_date": "2020-04-01T06:30:42.973000",
          "content": "<p>hi dude, i looked into yours. 2 points</p>\n\n<p><strong>1) memory overflow?</strong></p>\n\n<p>noticed that your way of generating 'frame1' consumed RAM upto 3.5GB with 400 files. this amount would be accumulated by the number of mp4 videos loaded, i'd say that it would exceed 4.9GB limit when you use it for private test set (4000 vids). so I, as other top rankers also disclosed, would like you to use ffmpeg.</p>\n\n<p><strong>2) generating submission.csv</strong></p>\n\n<p>since my way of writing codes are largely different from yours, let me provide simpler way in generating submission.csv as an example. try to replace your with this kind below.</p>\n\n<pre><code># predict with pretrained tf model first, and store predictions in a dict named 'result'\n    for mp4_filename in mp4_filename_list:\n        preds = model.predict_proba(x)\n        result[mp4_filename] = np.squeeze(preds[0][0])\n\n# generate submission.csv next.\nsample_submission = pd.read_csv('../input/deepfake-detection-challenge/sample_submission.csv')\nfor c in sample_submission.columns[sample_submission.columns == 'label']:\n    try:\n        sample_submission[c] = result.values()    # writing corresponding predicted value here\n    except:\n        sample_submission[c] = 0.5    # if no predicted value found, allocate 0.5 \nsample_submission.to_csv('submission.csv', index=False)\n</code></pre>\n\n<p>Nice if these above works for you even a bit. Let's see in the next competition again. :-)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 793819,
          "author_name": "Moshel",
          "author_url": "",
          "post_date": "2020-04-01T10:11:32.490000",
          "content": "<p>You should read carefully the competition getting started tab. Your code has 400 hard coded. When you submit, the test videos dir is replaced with a new one that has 4000 entries </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 793820,
          "author_name": "Moshel",
          "author_url": "",
          "post_date": "2020-04-01T10:12:14.810000",
          "content": "<p>Also, you should add proper error handling. Some of the 4000 videos are corrupted </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 794029,
          "author_name": "Debopriyo Sanyal",
          "author_url": "",
          "post_date": "2020-04-01T13:53:34.890000",
          "content": "<p>@Moshel Thanks for your observations.After you mentioned I checked the getting started tab again: 1. I did not see a mention of 4000 entries anywhere. 2. Regardless; I agree the coding should have been so done so as to handle any number of files and not just 400.3. I should have put in a condition to handle corrupted files, this just did not occur to me.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 790981,
      "author_name": "Brad Davis",
      "author_url": "",
      "post_date": "2020-03-30T02:26:54.003000",
      "content": "<p>I am getting a submission error as well - does this take away from my available submissions?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 790500,
      "author_name": "Akash",
      "author_url": "",
      "post_date": "2020-03-29T16:24:55.790000",
      "content": "<p>Make sure you do not save intermediate files to disk, kernel has 5GB of storage for output/working dir.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 790086,
      "author_name": "Vadasz",
      "author_url": "",
      "post_date": "2020-03-29T09:24:28.723000",
      "content": "<p>I had this error in different 2 cases:\n- I forgot the csv header, I wrote only the results.\n- I had nan value at some results. Check zero division, etc.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 790138,
          "author_name": "Debopriyo Sanyal",
          "author_url": "",
          "post_date": "2020-03-29T10:29:36.533000",
          "content": "<p>I am using  df.to_csv(\"submission.csv\", encoding='utf-8', index=False) to print my output to file. Would doing that create a csv header in the generated file? My understanding is that it does, however you know correct me if I am ...</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 790032,
      "author_name": "Mats",
      "author_url": "",
      "post_date": "2020-03-29T08:09:29.377000",
      "content": "<p>ds, I'm also suffering from same challenge (another topic was made just now). How much time does your notebook spend to run with 400 public test videos? I'm asking this because I heard that roughly 10 times of the time spent for public test data needs to be below 9 hours. Mine does 40 mins. If anything rings a bell with you, pls try check it out.   Mats</p>",
      "votes": 0,
      "replies": [
        {
          "id": 790050,
          "author_name": "Debopriyo Sanyal",
          "author_url": "",
          "post_date": "2020-03-29T08:31:55.487000",
          "content": "<p>my notebook when run cell by cell takes less than 30 mins.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 790058,
          "author_name": "Mats",
          "author_url": "",
          "post_date": "2020-03-29T08:41:18.437000",
          "content": "<p>ok. got it. just minutes ago, one smart person suggested me to output NOT '1, 0, 0, 1, ...' kind of value to submission.csv but to do actual float value. Mie was done only with integers. If your submission.csv also contains such integer only, worth trying it if you would like to. I'll keep following this topic.  I'm running my code now. Mats</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 790124,
          "author_name": "Debopriyo Sanyal",
          "author_url": "",
          "post_date": "2020-03-29T10:08:45.420000",
          "content": "<p>Did the float thing help solve your issue?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 790132,
          "author_name": "Mats",
          "author_url": "",
          "post_date": "2020-03-29T10:22:33.980000",
          "content": "<p>still working on it. will keep u posted when done.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 790972,
          "author_name": "Mats",
          "author_url": "",
          "post_date": "2020-03-30T02:02:28.207000",
          "content": "<p>hey buddy, float didn't work.</p>\n\n<p><a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/139528\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/139528</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 789946,
      "author_name": "Debopriyo Sanyal",
      "author_url": "",
      "post_date": "2020-03-29T06:30:37.007000",
      "content": "<p>Any ideas how to go about this submission scoring error? The code works okay when run in notebook.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 789983,
          "author_name": "Satish Yenumula",
          "author_url": "",
          "post_date": "2020-03-29T07:21:34.920000",
          "content": "<p>Could you specify what the error says or a screenshot?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 790023,
          "author_name": "Debopriyo Sanyal",
          "author_url": "",
          "post_date": "2020-03-29T08:03:17.857000",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3369259%2F1346a7aa524460851f09a694865653d2%2FScreenshot%202020-03-29%20at%201.32.19%20PM.png?generation=1585468990178305&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 790024,
          "author_name": "Debopriyo Sanyal",
          "author_url": "",
          "post_date": "2020-03-29T08:03:44.760000",
          "content": "<p>Please find above the error screenshot</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 792597,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-31T11:04:46.350000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "794449": "I've got submission scoring errors on notebooks that ran fine before, maybe that was because their run on public lb has finished past the deadline (though I've submitted those a few hours before it) Will have to wait for late submission to open to check it.",
    "793518": "Can someone tell what went wrong here:\n\nhttps://www.kaggle.com/ds123456/nuclear-zone?scriptVersionId=31176407",
    "790981": "I am getting a submission error as well - does this take away from my available submissions?",
    "790500": "Make sure you do not save intermediate files to disk, kernel has 5GB of storage for output/working dir.",
    "790086": "I had this error in different 2 cases:\n- I forgot the csv header, I wrote only the results.\n- I had nan value at some results. Check zero division, etc.",
    "790032": "ds, I'm also suffering from same challenge (another topic was made just now). How much time does your notebook spend to run with 400 public test videos? I'm asking this because I heard that roughly 10 times of the time spent for public test data needs to be below 9 hours. Mine does 40 mins. If anything rings a bell with you, pls try check it out.   Mats",
    "789946": "Any ideas how to go about this submission scoring error? The code works okay when run in notebook.",
    "789942": "I am getting a submission scoring error, the code works okay but whenever I submit it gives me a scoring error.",
    "792597": ""
  }
}