{
  "id": 39571,
  "title": "ERROR: File appears to be empty (Line 1, Column 1)",
  "url": "/competitions/carvana-image-masking-challenge/discussion/39571",
  "author_name": "",
  "post_date": "2017-09-16T17:48:04.976045800Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I'm encountering an error after uploading my submission - getting\n\"ERROR: File appears to be empty (Line 1, Column 1)\"</p>\n\n<p>My CSV is generated with Pandas, and I'm able to read it just fine, and decode the RLEs.\nThe CSV is then 7Zipped.</p>\n\n<p>One odd thing I noticed is that my CSV is 3GB while some users reported submission of ~1GB.</p>\n\n<p>Any ideas?</p>",
  "messages": [
    {
      "id": "221896",
      "postDate": "09/16/2017 17:48:04",
      "content": "<p>I'm encountering an error after uploading my submission - getting\n\"ERROR: File appears to be empty (Line 1, Column 1)\"</p>\n\n<p>My CSV is generated with Pandas, and I'm able to read it just fine, and decode the RLEs.\nThe CSV is then 7Zipped.</p>\n\n<p>One odd thing I noticed is that my CSV is 3GB while some users reported submission of ~1GB.</p>\n\n<p>Any ideas?</p>",
      "rawMarkdown": "I'm encountering an error after uploading my submission - getting\n\"ERROR: File appears to be empty (Line 1, Column 1)\"\n\nMy CSV is generated with Pandas, and I'm able to read it just fine, and decode the RLEs.\nThe CSV is then 7Zipped.\n\nOne odd thing I noticed is that my CSV is 3GB while some users reported submission of ~1GB.\n\nAny ideas?",
      "votes": null
    },
    {
      "id": "221960",
      "postDate": "09/17/2017 02:40:54",
      "content": "<p>As for your second question:\nIf your CSV file is that big you probably messed up your predictions somehow. Since it is RLE you would expect to encode a valid mask by not much more than 1280*2 numbers. This is because in most cases you would expect the intersection of the mask with a horizontal line to yield just one connected line again. This would yield approximately 10^5*2560 numbers if we multiplied that by an average of 4 bytes for the string representation of those numbers you get to approximately 1 GB filesize. Having a file three times the size of that means that your masks is highly unconnected. </p>",
      "rawMarkdown": "As for your second question:\nIf your CSV file is that big you probably messed up your predictions somehow. Since it is RLE you would expect to encode a valid mask by not much more than 1280*2 numbers. This is because in most cases you would expect the intersection of the mask with a horizontal line to yield just one connected line again. This would yield approximately 10^5*2560 numbers if we multiplied that by an average of 4 bytes for the string representation of those numbers you get to approximately 1 GB filesize. Having a file three times the size of that means that your masks is highly unconnected.",
      "votes": null
    },
    {
      "id": "221970",
      "postDate": "09/17/2017 04:25:12",
      "content": "<p>I've encountered the same case (3GB csv files). It took 3 days for me to find out that it's cause by the <em>BAD PREDICTED MASK</em> . More generally, unclear edges of those car masks.</p>\n\n<p>Noticed that if you got a predicted mask like this:</p>\n\n<p>0 1 0 1 0 1 0 1\n1 0 1 0 1 0 1 0\n0 1 0 1 0 1 0 1\n1 0 1 0 1 0 1 0</p>\n\n<p>it's easy to imagine that you will get an extremely large RLE sequence in this case.</p>",
      "rawMarkdown": "I've encountered the same case (3GB csv files). It took 3 days for me to find out that it's cause by the _BAD PREDICTED MASK_ . More generally, unclear edges of those car masks.\n\nNoticed that if you got a predicted mask like this:\n\n0 1 0 1 0 1 0 1\n1 0 1 0 1 0 1 0\n0 1 0 1 0 1 0 1\n1 0 1 0 1 0 1 0\n\nit's easy to imagine that you will get an extremely large RLE sequence in this case.",
      "votes": null
    },
    {
      "id": "221971",
      "postDate": "09/17/2017 04:31:35",
      "content": "<p>In this dataset, many ground truth masks got numbers of <em>holes</em> within the car, especially for the view ID 05 and 13. So it's hard to say that the RLE sequence is not more than 1280*2 numbers. </p>",
      "rawMarkdown": "In this dataset, many ground truth masks got numbers of _holes_ within the car, especially for the view ID 05 and 13. So it's hard to say that the RLE sequence is not more than 1280*2 numbers.",
      "votes": null
    },
    {
      "id": "221981",
      "postDate": "09/17/2017 05:54:16",
      "content": "<p>I've been able to fix the error in my masks - I had a bug causing my RLE masks to contain the entire batch.\nNow I'm getting another error \"Submission timed out. Please reduce size before resubmitting.\n\"</p>",
      "rawMarkdown": "I've been able to fix the error in my masks - I had a bug causing my RLE masks to contain the entire batch.\nNow I'm getting another error \"Submission timed out. Please reduce size before resubmitting.\n\"",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 221960,
      "author_name": "felixdo",
      "author_url": "",
      "post_date": "09/17/2017 02:40:54",
      "content": "<p>As for your second question:\nIf your CSV file is that big you probably messed up your predictions somehow. Since it is RLE you would expect to encode a valid mask by not much more than 1280*2 numbers. This is because in most cases you would expect the intersection of the mask with a horizontal line to yield just one connected line again. This would yield approximately 10^5*2560 numbers if we multiplied that by an average of 4 bytes for the string representation of those numbers you get to approximately 1 GB filesize. Having a file three times the size of that means that your masks is highly unconnected. </p>",
      "votes": null,
      "replies": [
        {
          "id": 221971,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "09/17/2017 04:31:35",
          "content": "<p>In this dataset, many ground truth masks got numbers of <em>holes</em> within the car, especially for the view ID 05 and 13. So it's hard to say that the RLE sequence is not more than 1280*2 numbers. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 221981,
          "author_name": "burgalon",
          "author_url": "",
          "post_date": "09/17/2017 05:54:16",
          "content": "<p>I've been able to fix the error in my masks - I had a bug causing my RLE masks to contain the entire batch.\nNow I'm getting another error \"Submission timed out. Please reduce size before resubmitting.\n\"</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 221970,
      "author_name": "haqishen",
      "author_url": "",
      "post_date": "09/17/2017 04:25:12",
      "content": "<p>I've encountered the same case (3GB csv files). It took 3 days for me to find out that it's cause by the <em>BAD PREDICTED MASK</em> . More generally, unclear edges of those car masks.</p>\n\n<p>Noticed that if you got a predicted mask like this:</p>\n\n<p>0 1 0 1 0 1 0 1\n1 0 1 0 1 0 1 0\n0 1 0 1 0 1 0 1\n1 0 1 0 1 0 1 0</p>\n\n<p>it's easy to imagine that you will get an extremely large RLE sequence in this case.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "221896": "I'm encountering an error after uploading my submission - getting\n\"ERROR: File appears to be empty (Line 1, Column 1)\"\n\nMy CSV is generated with Pandas, and I'm able to read it just fine, and decode the RLEs.\nThe CSV is then 7Zipped.\n\nOne odd thing I noticed is that my CSV is 3GB while some users reported submission of ~1GB.\n\nAny ideas?",
    "221960": "As for your second question:\nIf your CSV file is that big you probably messed up your predictions somehow. Since it is RLE you would expect to encode a valid mask by not much more than 1280*2 numbers. This is because in most cases you would expect the intersection of the mask with a horizontal line to yield just one connected line again. This would yield approximately 10^5*2560 numbers if we multiplied that by an average of 4 bytes for the string representation of those numbers you get to approximately 1 GB filesize. Having a file three times the size of that means that your masks is highly unconnected.",
    "221970": "I've encountered the same case (3GB csv files). It took 3 days for me to find out that it's cause by the _BAD PREDICTED MASK_ . More generally, unclear edges of those car masks.\n\nNoticed that if you got a predicted mask like this:\n\n0 1 0 1 0 1 0 1\n1 0 1 0 1 0 1 0\n0 1 0 1 0 1 0 1\n1 0 1 0 1 0 1 0\n\nit's easy to imagine that you will get an extremely large RLE sequence in this case.",
    "221971": "In this dataset, many ground truth masks got numbers of _holes_ within the car, especially for the view ID 05 and 13. So it's hard to say that the RLE sequence is not more than 1280*2 numbers.",
    "221981": "I've been able to fix the error in my masks - I had a bug causing my RLE masks to contain the entire batch.\nNow I'm getting another error \"Submission timed out. Please reduce size before resubmitting.\n\""
  },
  "source": "meta"
}