{
  "id": 3327,
  "title": "End of file problem",
  "url": "/competitions/flight/discussion/3327",
  "author_name": "",
  "post_date": "2012-12-14T12:23:03.337Z",
  "votes": null,
  "comment_count": 6,
  "views": 4449,
  "content": "<p>When I try to read the training data into my c&#43;&#43; program, I'm running into a problem at the end of file. &nbsp;For the InitialTrainingSet 2012_11_12 flighthistory.csv, the actual data (including '\\n' and spaces) is exactly 7784526 characters long. &nbsp;However, if\r\n I keep reading until I hit an end of file character (== '\\0'), I end up reading 7809940 characters, where the extra characters are gibberish. &nbsp;</p>\r\n<p>My text editor, too, insists that the file is 7809940 characters large, but a direct character count yields only 7784526 characters.</p>\r\n<p>Anyone know what's going on?</p>",
  "messages": [
    {
      "id": "17822",
      "postDate": "12/14/2012 12:23:03",
      "content": "<p>When I try to read the training data into my c&#43;&#43; program, I'm running into a problem at the end of file. &nbsp;For the InitialTrainingSet 2012_11_12 flighthistory.csv, the actual data (including '\\n' and spaces) is exactly 7784526 characters long. &nbsp;However, if\r\n I keep reading until I hit an end of file character (== '\\0'), I end up reading 7809940 characters, where the extra characters are gibberish. &nbsp;</p>\r\n<p>My text editor, too, insists that the file is 7809940 characters large, but a direct character count yields only 7784526 characters.</p>\r\n<p>Anyone know what's going on?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "17824",
      "postDate": "12/14/2012 13:48:01",
      "content": "<p>You are mistaken. Your &quot;direct character count&quot; didn't take into account new lines (0x0D 0x0A) - that is it didn't take into account 0x0D (\\r).</p>\r\n<p>Total number of bytes is&nbsp;<span>7809940.</span></p>\r\n<p>Total number of bytes excluding 0x0D is&nbsp;<span>7784526.</span></p>\r\n<p><span>You can also deduce that your &quot;direct character count&quot; didn't take into account new lines by subtracting&nbsp;7784526 from the total number of bytes&nbsp;7809940. You will get 25414 which corresponds to the number of lines in that file.</span></p>\r\n<p>This new line format (0x0D 0x0A) is typical DOS/Windows format - 0x0A is typical Linux format (<a href=\"https://en.wikipedia.org/wiki/Newline\">https://en.wikipedia.org/wiki/Newline</a>).</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "17847",
      "postDate": "12/14/2012 21:17:59",
      "content": "<p>OK, the plot thickens ...</p>\r\n<p>Thanks for pointing out the '\\r'. &nbsp;If I include this as a character as well, then you're right, &nbsp;the full size in the text editor is correct.</p>\r\n<p>However, when actually reading the data in (using fread) the '\\r' character seems to be completely ignored, it doesn't appear in the charcter stream that I get back. &nbsp;I suspect this is the issue.</p>\r\n<p>Here's how I'm doing the input:</p>\r\n<p>&nbsp;</p>\r\n<p>char* buffer = new char[large_size];</p>\r\n<p>fread(buffer,1,large_size,file);</p>\r\n<p>&nbsp;</p>\r\n<p>As I read through the data, because the '\\r' are not present in the stream, I finish reading the last entry after&nbsp;<span>7784526 characters. &nbsp;</span></p>\r\n<p><span>Then a weird thing happens.&nbsp;</span></p>\r\n<p><span>If I query the buffer again, it jumps back in the file, exactly&nbsp;<span>25414 characters and continues spitting out the same data again until I hit the end of line.</span></span></p>\r\n<p>This jump is exactly the number of missing '\\r' characters, so I rather suspect this is the culprit.</p>\r\n<p>&nbsp;</p>\r\n<p>Who can solve the mystery of the missing carriage returns?!</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "17848",
      "postDate": "12/14/2012 21:29:25",
      "content": "<p>Try Python.</p>\r\n<p>&nbsp;</p>\r\n<p>These kinds of problems are usually best solved in productivity focused languages with little concern for performance optimization. The underlying algorithms you design should take into consideration the nature of the problem and provide for scalability\r\n across a large dataset.</p>\r\n<p>&nbsp;</p>\r\n<p>After you have validated your algorithm, then you might want to start optimizing for performance.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "17851",
      "postDate": "12/14/2012 22:11:04",
      "content": "<p>One of these days I'll get around to learning Python ...</p>\r\n<p>&nbsp;</p>\r\n<p>If anyone else is having this same problem, I've found a work-around. &nbsp;fread returns the number of characters read, so I don't have to go hunting for an end of file character. &nbsp;For the data set I've been using, fread returns the smaller number&nbsp;<span>7784526.\r\n &nbsp;I have no idea why the end of file is located far beyond this point.</span></p>\r\n<p><span><br>\r\n</span></p>\r\n<p>I'm still curious about what is causing this issue, if anyone can enlighten me ...</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "17852",
      "postDate": "12/14/2012 23:10:54",
      "content": "<p>Are you checking the file pointer for EOF correctly?</p>\r\n<p>e.g. &nbsp; &nbsp;while(!feof(fp)) { /* do stuff */ }</p>\r\n<p>I always managed to mangle that part somehow (e.g. using input == EOF)</p>\r\n<p>Also, this may be dependant on what OS/compiler combination you're using. I seem to remember some of them try to be fancy (read: non-standard) and deal with linux/windows/mac line ending conversions, others don't. Although it's been a very long time since\r\n I've had to deal with low level file IO.</p>\r\n<p>Seriously though, C/C&#43;&#43; would be [almost] the last language I would use to do this. That you're having this problem seems proof enough.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "17903",
      "postDate": "12/16/2012 03:25:32",
      "content": "<p>I'm actually looking at the character stream and stopping when I see the '\\0' character. &nbsp;Perhaps feof would work properly.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 17824,
      "author_name": "predragbradaric",
      "author_url": "",
      "post_date": "12/14/2012 13:48:01",
      "content": "<p>You are mistaken. Your &quot;direct character count&quot; didn't take into account new lines (0x0D 0x0A) - that is it didn't take into account 0x0D (\\r).</p>\r\n<p>Total number of bytes is&nbsp;<span>7809940.</span></p>\r\n<p>Total number of bytes excluding 0x0D is&nbsp;<span>7784526.</span></p>\r\n<p><span>You can also deduce that your &quot;direct character count&quot; didn't take into account new lines by subtracting&nbsp;7784526 from the total number of bytes&nbsp;7809940. You will get 25414 which corresponds to the number of lines in that file.</span></p>\r\n<p>This new line format (0x0D 0x0A) is typical DOS/Windows format - 0x0A is typical Linux format (<a href=\"https://en.wikipedia.org/wiki/Newline\">https://en.wikipedia.org/wiki/Newline</a>).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 17847,
      "author_name": "lonelyatthetop",
      "author_url": "",
      "post_date": "12/14/2012 21:17:59",
      "content": "<p>OK, the plot thickens ...</p>\r\n<p>Thanks for pointing out the '\\r'. &nbsp;If I include this as a character as well, then you're right, &nbsp;the full size in the text editor is correct.</p>\r\n<p>However, when actually reading the data in (using fread) the '\\r' character seems to be completely ignored, it doesn't appear in the charcter stream that I get back. &nbsp;I suspect this is the issue.</p>\r\n<p>Here's how I'm doing the input:</p>\r\n<p>&nbsp;</p>\r\n<p>char* buffer = new char[large_size];</p>\r\n<p>fread(buffer,1,large_size,file);</p>\r\n<p>&nbsp;</p>\r\n<p>As I read through the data, because the '\\r' are not present in the stream, I finish reading the last entry after&nbsp;<span>7784526 characters. &nbsp;</span></p>\r\n<p><span>Then a weird thing happens.&nbsp;</span></p>\r\n<p><span>If I query the buffer again, it jumps back in the file, exactly&nbsp;<span>25414 characters and continues spitting out the same data again until I hit the end of line.</span></span></p>\r\n<p>This jump is exactly the number of missing '\\r' characters, so I rather suspect this is the culprit.</p>\r\n<p>&nbsp;</p>\r\n<p>Who can solve the mystery of the missing carriage returns?!</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 17848,
      "author_name": "rabeti",
      "author_url": "",
      "post_date": "12/14/2012 21:29:25",
      "content": "<p>Try Python.</p>\r\n<p>&nbsp;</p>\r\n<p>These kinds of problems are usually best solved in productivity focused languages with little concern for performance optimization. The underlying algorithms you design should take into consideration the nature of the problem and provide for scalability\r\n across a large dataset.</p>\r\n<p>&nbsp;</p>\r\n<p>After you have validated your algorithm, then you might want to start optimizing for performance.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 17851,
      "author_name": "lonelyatthetop",
      "author_url": "",
      "post_date": "12/14/2012 22:11:04",
      "content": "<p>One of these days I'll get around to learning Python ...</p>\r\n<p>&nbsp;</p>\r\n<p>If anyone else is having this same problem, I've found a work-around. &nbsp;fread returns the number of characters read, so I don't have to go hunting for an end of file character. &nbsp;For the data set I've been using, fread returns the smaller number&nbsp;<span>7784526.\r\n &nbsp;I have no idea why the end of file is located far beyond this point.</span></p>\r\n<p><span><br>\r\n</span></p>\r\n<p>I'm still curious about what is causing this issue, if anyone can enlighten me ...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 17852,
      "author_name": "cephalopod",
      "author_url": "",
      "post_date": "12/14/2012 23:10:54",
      "content": "<p>Are you checking the file pointer for EOF correctly?</p>\r\n<p>e.g. &nbsp; &nbsp;while(!feof(fp)) { /* do stuff */ }</p>\r\n<p>I always managed to mangle that part somehow (e.g. using input == EOF)</p>\r\n<p>Also, this may be dependant on what OS/compiler combination you're using. I seem to remember some of them try to be fancy (read: non-standard) and deal with linux/windows/mac line ending conversions, others don't. Although it's been a very long time since\r\n I've had to deal with low level file IO.</p>\r\n<p>Seriously though, C/C&#43;&#43; would be [almost] the last language I would use to do this. That you're having this problem seems proof enough.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 17903,
      "author_name": "lonelyatthetop",
      "author_url": "",
      "post_date": "12/16/2012 03:25:32",
      "content": "<p>I'm actually looking at the character stream and stopping when I see the '\\0' character. &nbsp;Perhaps feof would work properly.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "17822": "",
    "17824": "",
    "17847": "",
    "17848": "",
    "17851": "",
    "17852": "",
    "17903": ""
  },
  "source": "meta"
}