{
  "id": 27593,
  "title": "Why pandas read both strings and integers value for platforms in events.csv",
  "url": "/competitions/outbrain-click-prediction/discussion/27593",
  "author_name": "",
  "post_date": "2017-01-11T13:19:46.037Z",
  "votes": null,
  "comment_count": 3,
  "views": 92,
  "content": "<p>command:</p>\n\n<pre><code>events = pd.read_csv(eFile) #eFile is path of event file\nevents[\"platform\"].unique()\n</code></pre>\n\n<p>outputs:</p>\n\n<pre><code>array([3, 2, 1, '2', '1', '3', '\\\\N'], dtype=object)\n</code></pre>\n\n<p>There are five rows in the file which have platform == '\\N', so I understand the value '\\\\N'. But I am not able to understand why values, 1, 2 and 3 are being read sometimes as string and sometimes as integers?</p>\n\n<p>I thought, may be once read_csv encounters a string value ('\\N'), it switched to string, but I found thats not the case (Values '\\N' is encountered much later after read_csv reads 1, 2 and 3 as string.  I found that first display_id for string values are, as outputted are ['1': 262147, '2': 262145, '3': 262155, '\\N': 303066].\n(by using commands like: events[events[\"platform\"] == '2'].head())</p>\n\n<p>When I use cat command to see if there is any difference between platform column among rows for which pandas is reading values as string verses integers, I find no such difference. In both cases values are simply 1, 2 or 3.</p>\n\n<p>So why is read_csv behaving in this way? </p>\n\n<p>Note: After loading events file, I manually changed '\\N' values as 0 for platform. Then I wrote the DataFrame to csv file and tried to read it back. In this case I get only 4 values [0, 1, 2, 3], as expected.</p>",
  "messages": [
    {
      "id": "155479",
      "postDate": "01/11/2017 13:19:46",
      "content": "<p>command:</p>\n\n<pre><code>events = pd.read_csv(eFile) #eFile is path of event file\nevents[\"platform\"].unique()\n</code></pre>\n\n<p>outputs:</p>\n\n<pre><code>array([3, 2, 1, '2', '1', '3', '\\\\N'], dtype=object)\n</code></pre>\n\n<p>There are five rows in the file which have platform == '\\N', so I understand the value '\\\\N'. But I am not able to understand why values, 1, 2 and 3 are being read sometimes as string and sometimes as integers?</p>\n\n<p>I thought, may be once read_csv encounters a string value ('\\N'), it switched to string, but I found thats not the case (Values '\\N' is encountered much later after read_csv reads 1, 2 and 3 as string.  I found that first display_id for string values are, as outputted are ['1': 262147, '2': 262145, '3': 262155, '\\N': 303066].\n(by using commands like: events[events[\"platform\"] == '2'].head())</p>\n\n<p>When I use cat command to see if there is any difference between platform column among rows for which pandas is reading values as string verses integers, I find no such difference. In both cases values are simply 1, 2 or 3.</p>\n\n<p>So why is read_csv behaving in this way? </p>\n\n<p>Note: After loading events file, I manually changed '\\N' values as 0 for platform. Then I wrote the DataFrame to csv file and tried to read it back. In this case I get only 4 values [0, 1, 2, 3], as expected.</p>",
      "rawMarkdown": "command:\r\n\r\n    events = pd.read_csv(eFile) #eFile is path of event file\r\n    events[\"platform\"].unique()\r\n\r\noutputs:\r\n\r\n    array([3, 2, 1, '2', '1', '3', '\\\\N'], dtype=object)\r\n\r\nThere are five rows in the file which have platform == '\\\\N', so I understand the value '\\\\\\\\N'. But I am not able to understand why values, 1, 2 and 3 are being read sometimes as string and sometimes as integers?\r\n\r\nI thought, may be once read_csv encounters a string value ('\\\\N'), it switched to string, but I found thats not the case (Values '\\\\N' is encountered much later after read_csv reads 1, 2 and 3 as string.  I found that first display_id for string values are, as outputted are ['1': 262147, '2': 262145, '3': 262155, '\\\\N': 303066].\r\n(by using commands like: events[events[\"platform\"] == '2'].head())\r\n\r\nWhen I use cat command to see if there is any difference between platform column among rows for which pandas is reading values as string verses integers, I find no such difference. In both cases values are simply 1, 2 or 3.\r\n\r\nSo why is read_csv behaving in this way? \r\n\r\nNote: After loading events file, I manually changed '\\\\N' values as 0 for platform. Then I wrote the DataFrame to csv file and tried to read it back. In this case I get only 4 values [0, 1, 2, 3], as expected.",
      "votes": null
    },
    {
      "id": "155612",
      "postDate": "01/12/2017 01:24:30",
      "content": "<p>[quote=AshutoshNirala;155479]</p>\n\n<p>command:</p>\n\n<pre><code>events = pd.read_csv(eFile) #eFile is path of event file\nevents[\"platform\"].unique()\n</code></pre>\n\n<p>outputs:</p>\n\n<pre><code>array([3, 2, 1, '2', '1', '3', '\\\\N'], dtype=object)\n</code></pre>\n\n<p>There are five rows in the file which have platform == '\\N', so I understand the value '\\\\N'. But I am not able to understand why values, 1, 2 and 3 are being read sometimes as string and sometimes as integers?</p>\n\n<p>I thought, may be once read_csv encounters a string value ('\\N'), it switched to string, but I found thats not the case (Values '\\N' is encountered much later after read_csv reads 1, 2 and 3 as string.  I found that first display_id for string values are, as outputted are ['1': 262147, '2': 262145, '3': 262155, '\\N': 303066].\n(by using commands like: events[events[\"platform\"] == '2'].head())</p>\n\n<p>When I use cat command to see if there is any difference between platform column among rows for which pandas is reading values as string verses integers, I find no such difference. In both cases values are simply 1, 2 or 3.</p>\n\n<p>So why is read_csv behaving in this way? </p>\n\n<p>Note: After loading events file, I manually changed '\\N' values as 0 for platform. Then I wrote the DataFrame to csv file and tried to read it back. In this case I get only 4 values [0, 1, 2, 3], as expected.</p>\n\n<p>[/quote]\nThat's Outbrain's problem. They represented platforms as both int and string. What we need to do is to convert string to int.</p>",
      "rawMarkdown": "[quote=AshutoshNirala;155479]\r\n\r\ncommand:\r\n\r\n    events = pd.read_csv(eFile) #eFile is path of event file\r\n    events[\"platform\"].unique()\r\n\r\noutputs:\r\n\r\n    array([3, 2, 1, '2', '1', '3', '\\\\N'], dtype=object)\r\n\r\nThere are five rows in the file which have platform == '\\\\N', so I understand the value '\\\\\\\\N'. But I am not able to understand why values, 1, 2 and 3 are being read sometimes as string and sometimes as integers?\r\n\r\nI thought, may be once read_csv encounters a string value ('\\\\N'), it switched to string, but I found thats not the case (Values '\\\\N' is encountered much later after read_csv reads 1, 2 and 3 as string.  I found that first display_id for string values are, as outputted are ['1': 262147, '2': 262145, '3': 262155, '\\\\N': 303066].\r\n(by using commands like: events[events[\"platform\"] == '2'].head())\r\n\r\nWhen I use cat command to see if there is any difference between platform column among rows for which pandas is reading values as string verses integers, I find no such difference. In both cases values are simply 1, 2 or 3.\r\n\r\nSo why is read_csv behaving in this way? \r\n\r\nNote: After loading events file, I manually changed '\\\\N' values as 0 for platform. Then I wrote the DataFrame to csv file and tried to read it back. In this case I get only 4 values [0, 1, 2, 3], as expected.\r\n\r\n[/quote]\r\nThat's Outbrain's problem. They represented platforms as both int and string. What we need to do is to convert string to int.",
      "votes": null
    },
    {
      "id": "155653",
      "postDate": "01/12/2017 09:11:13",
      "content": "<p>Thanks @peixiang</p>\n\n<p>But, I can't find any difference between the representation for [platform] column of rows, in events.csv file, when it is being read as string and when it is being read as integers.</p>\n\n<p>For example: Row 1 (counting header as Row 0), of event.csv file, where platform value 3 is being read as int, is:</p>\n\n<p>1,cb8c55702adb93,379743,61,3,US&gt;SC&gt;519</p>\n\n<p>And row 262155, where platform value 3 is being read as string '3' is:</p>\n\n<p>262155,98b21ea98f96e2,1783753,26385152,3,US&gt;NY&gt;501</p>\n\n<p>Note that in both cases platform is represented as simply: ....,3,.... (no quotes or anything for string). So, on what basis is pandas reading value in one row string and another int?</p>",
      "rawMarkdown": "Thanks @peixiang\r\n\r\nBut, I can't find any difference between the representation for [platform] column of rows, in events.csv file, when it is being read as string and when it is being read as integers.\r\n\r\nFor example: Row 1 (counting header as Row 0), of event.csv file, where platform value 3 is being read as int, is:\r\n\r\n1,cb8c55702adb93,379743,61,3,US>SC>519\r\n\r\nAnd row 262155, where platform value 3 is being read as string '3' is:\r\n\r\n262155,98b21ea98f96e2,1783753,26385152,3,US>NY>501\r\n\r\nNote that in both cases platform is represented as simply: ....,3,.... (no quotes or anything for string). So, on what basis is pandas reading value in one row string and another int?",
      "votes": null
    },
    {
      "id": "155655",
      "postDate": "01/12/2017 09:26:09",
      "content": "<p>The problem is that CSV files don't store the data type of the columns and pandas tries to infer the data type from the data. I believe pandas does that by looking at a sample of the values. Now, if this sample contains a \"\\N\", it assumes type string. If there is no \"\\N\" in the sample it assumes these are integers. </p>\n\n<p>There are several options to improve the parsing. You can for example pass a dictionary dtypes to read_csv and tell pandas explicitly what types you want. Or you can just convert afterwards, as you already did. You can also tell pandas to interpret \"\\N\" as missing values by passing 'na_values=\"\\N\"' to read_csv.</p>",
      "rawMarkdown": "The problem is that CSV files don't store the data type of the columns and pandas tries to infer the data type from the data. I believe pandas does that by looking at a sample of the values. Now, if this sample contains a \"\\\\N\", it assumes type string. If there is no \"\\\\N\" in the sample it assumes these are integers. \r\n\r\nThere are several options to improve the parsing. You can for example pass a dictionary dtypes to read_csv and tell pandas explicitly what types you want. Or you can just convert afterwards, as you already did. You can also tell pandas to interpret \"\\\\N\" as missing values by passing 'na_values=\"\\\\N\"' to read_csv.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 155612,
      "author_name": "peixiang",
      "author_url": "",
      "post_date": "01/12/2017 01:24:30",
      "content": "<p>[quote=AshutoshNirala;155479]</p>\n\n<p>command:</p>\n\n<pre><code>events = pd.read_csv(eFile) #eFile is path of event file\nevents[\"platform\"].unique()\n</code></pre>\n\n<p>outputs:</p>\n\n<pre><code>array([3, 2, 1, '2', '1', '3', '\\\\N'], dtype=object)\n</code></pre>\n\n<p>There are five rows in the file which have platform == '\\N', so I understand the value '\\\\N'. But I am not able to understand why values, 1, 2 and 3 are being read sometimes as string and sometimes as integers?</p>\n\n<p>I thought, may be once read_csv encounters a string value ('\\N'), it switched to string, but I found thats not the case (Values '\\N' is encountered much later after read_csv reads 1, 2 and 3 as string.  I found that first display_id for string values are, as outputted are ['1': 262147, '2': 262145, '3': 262155, '\\N': 303066].\n(by using commands like: events[events[\"platform\"] == '2'].head())</p>\n\n<p>When I use cat command to see if there is any difference between platform column among rows for which pandas is reading values as string verses integers, I find no such difference. In both cases values are simply 1, 2 or 3.</p>\n\n<p>So why is read_csv behaving in this way? </p>\n\n<p>Note: After loading events file, I manually changed '\\N' values as 0 for platform. Then I wrote the DataFrame to csv file and tried to read it back. In this case I get only 4 values [0, 1, 2, 3], as expected.</p>\n\n<p>[/quote]\nThat's Outbrain's problem. They represented platforms as both int and string. What we need to do is to convert string to int.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 155653,
      "author_name": "aknirala",
      "author_url": "",
      "post_date": "01/12/2017 09:11:13",
      "content": "<p>Thanks @peixiang</p>\n\n<p>But, I can't find any difference between the representation for [platform] column of rows, in events.csv file, when it is being read as string and when it is being read as integers.</p>\n\n<p>For example: Row 1 (counting header as Row 0), of event.csv file, where platform value 3 is being read as int, is:</p>\n\n<p>1,cb8c55702adb93,379743,61,3,US&gt;SC&gt;519</p>\n\n<p>And row 262155, where platform value 3 is being read as string '3' is:</p>\n\n<p>262155,98b21ea98f96e2,1783753,26385152,3,US&gt;NY&gt;501</p>\n\n<p>Note that in both cases platform is represented as simply: ....,3,.... (no quotes or anything for string). So, on what basis is pandas reading value in one row string and another int?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 155655,
      "author_name": "steitz",
      "author_url": "",
      "post_date": "01/12/2017 09:26:09",
      "content": "<p>The problem is that CSV files don't store the data type of the columns and pandas tries to infer the data type from the data. I believe pandas does that by looking at a sample of the values. Now, if this sample contains a \"\\N\", it assumes type string. If there is no \"\\N\" in the sample it assumes these are integers. </p>\n\n<p>There are several options to improve the parsing. You can for example pass a dictionary dtypes to read_csv and tell pandas explicitly what types you want. Or you can just convert afterwards, as you already did. You can also tell pandas to interpret \"\\N\" as missing values by passing 'na_values=\"\\N\"' to read_csv.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "155479": "command:\r\n\r\n    events = pd.read_csv(eFile) #eFile is path of event file\r\n    events[\"platform\"].unique()\r\n\r\noutputs:\r\n\r\n    array([3, 2, 1, '2', '1', '3', '\\\\N'], dtype=object)\r\n\r\nThere are five rows in the file which have platform == '\\\\N', so I understand the value '\\\\\\\\N'. But I am not able to understand why values, 1, 2 and 3 are being read sometimes as string and sometimes as integers?\r\n\r\nI thought, may be once read_csv encounters a string value ('\\\\N'), it switched to string, but I found thats not the case (Values '\\\\N' is encountered much later after read_csv reads 1, 2 and 3 as string.  I found that first display_id for string values are, as outputted are ['1': 262147, '2': 262145, '3': 262155, '\\\\N': 303066].\r\n(by using commands like: events[events[\"platform\"] == '2'].head())\r\n\r\nWhen I use cat command to see if there is any difference between platform column among rows for which pandas is reading values as string verses integers, I find no such difference. In both cases values are simply 1, 2 or 3.\r\n\r\nSo why is read_csv behaving in this way? \r\n\r\nNote: After loading events file, I manually changed '\\\\N' values as 0 for platform. Then I wrote the DataFrame to csv file and tried to read it back. In this case I get only 4 values [0, 1, 2, 3], as expected.",
    "155612": "[quote=AshutoshNirala;155479]\r\n\r\ncommand:\r\n\r\n    events = pd.read_csv(eFile) #eFile is path of event file\r\n    events[\"platform\"].unique()\r\n\r\noutputs:\r\n\r\n    array([3, 2, 1, '2', '1', '3', '\\\\N'], dtype=object)\r\n\r\nThere are five rows in the file which have platform == '\\\\N', so I understand the value '\\\\\\\\N'. But I am not able to understand why values, 1, 2 and 3 are being read sometimes as string and sometimes as integers?\r\n\r\nI thought, may be once read_csv encounters a string value ('\\\\N'), it switched to string, but I found thats not the case (Values '\\\\N' is encountered much later after read_csv reads 1, 2 and 3 as string.  I found that first display_id for string values are, as outputted are ['1': 262147, '2': 262145, '3': 262155, '\\\\N': 303066].\r\n(by using commands like: events[events[\"platform\"] == '2'].head())\r\n\r\nWhen I use cat command to see if there is any difference between platform column among rows for which pandas is reading values as string verses integers, I find no such difference. In both cases values are simply 1, 2 or 3.\r\n\r\nSo why is read_csv behaving in this way? \r\n\r\nNote: After loading events file, I manually changed '\\\\N' values as 0 for platform. Then I wrote the DataFrame to csv file and tried to read it back. In this case I get only 4 values [0, 1, 2, 3], as expected.\r\n\r\n[/quote]\r\nThat's Outbrain's problem. They represented platforms as both int and string. What we need to do is to convert string to int.",
    "155653": "Thanks @peixiang\r\n\r\nBut, I can't find any difference between the representation for [platform] column of rows, in events.csv file, when it is being read as string and when it is being read as integers.\r\n\r\nFor example: Row 1 (counting header as Row 0), of event.csv file, where platform value 3 is being read as int, is:\r\n\r\n1,cb8c55702adb93,379743,61,3,US>SC>519\r\n\r\nAnd row 262155, where platform value 3 is being read as string '3' is:\r\n\r\n262155,98b21ea98f96e2,1783753,26385152,3,US>NY>501\r\n\r\nNote that in both cases platform is represented as simply: ....,3,.... (no quotes or anything for string). So, on what basis is pandas reading value in one row string and another int?",
    "155655": "The problem is that CSV files don't store the data type of the columns and pandas tries to infer the data type from the data. I believe pandas does that by looking at a sample of the values. Now, if this sample contains a \"\\\\N\", it assumes type string. If there is no \"\\\\N\" in the sample it assumes these are integers. \r\n\r\nThere are several options to improve the parsing. You can for example pass a dictionary dtypes to read_csv and tell pandas explicitly what types you want. Or you can just convert afterwards, as you already did. You can also tell pandas to interpret \"\\\\N\" as missing values by passing 'na_values=\"\\\\N\"' to read_csv."
  },
  "source": "meta"
}