{
  "id": 21306,
  "title": "What is the encoding of ItemInfo_train data",
  "url": "/competitions/avito-duplicate-ads-detection/discussion/21306",
  "author_name": "",
  "post_date": "2016-05-30T03:23:22.070Z",
  "votes": null,
  "comment_count": 6,
  "views": 615,
  "content": "<p>I am having difficulty of reading some fields of this file, namely, description, title and attrsJSON. Is there anybody who can provide some script on reading this fields properly?</p>",
  "messages": [
    {
      "id": "121811",
      "postDate": "05/30/2016 03:23:22",
      "content": "<p>I am having difficulty of reading some fields of this file, namely, description, title and attrsJSON. Is there anybody who can provide some script on reading this fields properly?</p>",
      "rawMarkdown": "I am having difficulty of reading some fields of this file, namely, description, title and attrsJSON. Is there anybody who can provide some script on reading this fields properly?",
      "votes": null
    },
    {
      "id": "121821",
      "postDate": "05/30/2016 06:19:33",
      "content": "<p>In python:</p>\n\n<pre><code>import pandas as pd\ninfo = pd.read_csv( &quot;../input/ItemInfo_train.csv&quot;, encoding=&quot;utf-8&quot;)\n</code></pre>\n\n<p>For R, it may depend on the package you're using. I see solutions where they either pass it as a parameter (encoding='utf-8') or they change the locale for the script:</p>\n\n<pre><code>Sys.setlocale(category=&quot;LC_ALL&quot;, locale = &quot;ru_RU&quot;)\n</code></pre>\n\n<p>or:</p>\n\n<pre><code>df_ch &lt;- read.table(&quot;../input/ItemInfo_train.csv&quot;,\n                     sep=&quot;,&quot;,\n                     header=FALSE, \n                     encoding=&quot;russian&quot;, \n                     stringsAsFactors=FALSE\n                    )\n</code></pre>\n\n<p>Someone better experienced with R can probably give you a hint on which package should be used, they usually have different performances.</p>",
      "rawMarkdown": "In python:\r\n\r\n    import pandas as pd\r\n    info = pd.read_csv( \"../input/ItemInfo_train.csv\", encoding=\"utf-8\")\r\n\r\nFor R, it may depend on the package you're using. I see solutions where they either pass it as a parameter (encoding='utf-8') or they change the locale for the script:\r\n\r\n    Sys.setlocale(category=\"LC_ALL\", locale = \"ru_RU\")\r\n\r\nor:\r\n\r\n    df_ch <- read.table(\"../input/ItemInfo_train.csv\",\r\n                         sep=\",\",\r\n                         header=FALSE, \r\n                         encoding=\"russian\", \r\n                         stringsAsFactors=FALSE\r\n                        )\r\n\r\nSomeone better experienced with R can probably give you a hint on which package should be used, they usually have different performances.",
      "votes": null
    },
    {
      "id": "121828",
      "postDate": "05/30/2016 08:18:40",
      "content": "<pre><code>import pandas as pd\ninfo = pd.read_csv( &quot;../input/ItemInfo_train.csv&quot;, encoding=&quot;utf-8&quot;)\n</code></pre>\n\n<p>But this is taking too much time to load into memory, almost one minute, why is it extremely slow?</p>",
      "rawMarkdown": "import pandas as pd\r\n    info = pd.read_csv( \"../input/ItemInfo_train.csv\", encoding=\"utf-8\")\r\n\r\nBut this is taking too much time to load into memory, almost one minute, why is it extremely slow?",
      "votes": null
    },
    {
      "id": "121829",
      "postDate": "05/30/2016 08:21:51",
      "content": "<p>@humoyun, have a look at the file size</p>",
      "rawMarkdown": "humoyun, have a look at the file size",
      "votes": null
    },
    {
      "id": "121830",
      "postDate": "05/30/2016 08:30:02",
      "content": "<pre><code>import csv\nimport json\n\nfilepath='D:/DATAFOLDER/kaggle_avito/ItemInfo_train.csv'\ncount=0\n\nwith open(filepath) as csvfile:\n    readCSV = csv.reader(csvfile, delimiter=',')\n    for row in readCSV:\n        # tmp = list(row)\n        # fmt=u'{:&lt;15}'*len(tmp)\n        # print fmt.format(*[s.decode('utf-8') for s in tmp])\n        print(&quot;[{0}] (0 row)  itemID       : {1} &quot;.format(count, row[0]))\n        print(&quot;[{0}] (1 row)  categoryID   : {1} &quot;.format(count, row[1]))\n        print(&quot;[{0}] (2 row)  title        : {1} &quot;.format(count, unicode(row[2])))\n        print(&quot;[{0}] (3 row)  description  : {1} &quot;.format(count, unicode(row[3])))\n        print(&quot;[{0}] (4 row)  images_array : {1} &quot;.format(count, row[4]))\n        print(&quot;[{0}] (5 row)  attrsJSON    : {1} &quot;.format(count, unicode(row[5])))\n        #data = json.loads(row[5])\n        print(&quot;[{0}] (6 row)  price        : {1} &quot;.format(count, row[6]))\n        print(&quot;[{0}] (7 row)  locationID   : {1} &quot;.format(count, row[7]))\n        print(&quot;[{0}] (8 row)  metroID      : {1} &quot;.format(count, row[8]))\n        print(&quot;[{0}] (9 row)  lat          : {1} &quot;.format(count, row[9]))\n        print(&quot;[{0}] (10 row) long         : {1} &quot;.format(count, row[10]))\n\n        print &quot;**********************************************************************************&quot;\n        count+=1\n</code></pre>\n\n<p>I was trying to use this script to explore data set, but It gives me <strong>UnicodeDecodeError</strong>: ascii codec can't decode byte 0xd0 in position...</p>\n\n<p>@ololo  Can you tweak my script so that I can read the file properly</p>",
      "rawMarkdown": "import csv\r\n    import json\r\n    \r\n    filepath='D:/DATAFOLDER/kaggle_avito/ItemInfo_train.csv'\r\n    count=0\r\n    \r\n    with open(filepath) as csvfile:\r\n        readCSV = csv.reader(csvfile, delimiter=',')\r\n        for row in readCSV:\r\n        \t# tmp = list(row)\r\n        \t# fmt=u'{:<15}'*len(tmp)\r\n        \t# print fmt.format(*[s.decode('utf-8') for s in tmp])\r\n            print(\"[{0}] (0 row)  itemID \t   : {1} \".format(count, row[0]))\r\n            print(\"[{0}] (1 row)  categoryID   : {1} \".format(count, row[1]))\r\n            print(\"[{0}] (2 row)  title \t   : {1} \".format(count, unicode(row[2])))\r\n            print(\"[{0}] (3 row)  description  : {1} \".format(count, unicode(row[3])))\r\n            print(\"[{0}] (4 row)  images_array : {1} \".format(count, row[4]))\r\n            print(\"[{0}] (5 row)  attrsJSON    : {1} \".format(count, unicode(row[5])))\r\n            #data = json.loads(row[5])\r\n            print(\"[{0}] (6 row)  price        : {1} \".format(count, row[6]))\r\n            print(\"[{0}] (7 row)  locationID   : {1} \".format(count, row[7]))\r\n            print(\"[{0}] (8 row)  metroID  \t   : {1} \".format(count, row[8]))\r\n            print(\"[{0}] (9 row)  lat  \t\t   : {1} \".format(count, row[9]))\r\n            print(\"[{0}] (10 row) long         : {1} \".format(count, row[10]))\r\n         \r\n            print \"**********************************************************************************\"\r\n            count+=1\r\n\r\nI was trying to use this script to explore data set, but It gives me **UnicodeDecodeError**: ascii codec can't decode byte 0xd0 in position...\r\n\r\n@ololo  Can you tweak my script so that I can read the file properly\r\n\r\n\r\n\r\n\r\n  [1]: http://",
      "votes": null
    },
    {
      "id": "121833",
      "postDate": "05/30/2016 08:37:51",
      "content": "<p>use codecs.open instead </p>",
      "rawMarkdown": "use codecs.open instead",
      "votes": null
    },
    {
      "id": "121836",
      "postDate": "05/30/2016 08:46:01",
      "content": "<p>@ololo same output</p>",
      "rawMarkdown": "ololo same output",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 121821,
      "author_name": "remap1",
      "author_url": "",
      "post_date": "05/30/2016 06:19:33",
      "content": "<p>In python:</p>\n\n<pre><code>import pandas as pd\ninfo = pd.read_csv( &quot;../input/ItemInfo_train.csv&quot;, encoding=&quot;utf-8&quot;)\n</code></pre>\n\n<p>For R, it may depend on the package you're using. I see solutions where they either pass it as a parameter (encoding='utf-8') or they change the locale for the script:</p>\n\n<pre><code>Sys.setlocale(category=&quot;LC_ALL&quot;, locale = &quot;ru_RU&quot;)\n</code></pre>\n\n<p>or:</p>\n\n<pre><code>df_ch &lt;- read.table(&quot;../input/ItemInfo_train.csv&quot;,\n                     sep=&quot;,&quot;,\n                     header=FALSE, \n                     encoding=&quot;russian&quot;, \n                     stringsAsFactors=FALSE\n                    )\n</code></pre>\n\n<p>Someone better experienced with R can probably give you a hint on which package should be used, they usually have different performances.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121828,
      "author_name": "humoyun",
      "author_url": "",
      "post_date": "05/30/2016 08:18:40",
      "content": "<pre><code>import pandas as pd\ninfo = pd.read_csv( &quot;../input/ItemInfo_train.csv&quot;, encoding=&quot;utf-8&quot;)\n</code></pre>\n\n<p>But this is taking too much time to load into memory, almost one minute, why is it extremely slow?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121829,
      "author_name": "agrigorev",
      "author_url": "",
      "post_date": "05/30/2016 08:21:51",
      "content": "<p>@humoyun, have a look at the file size</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121830,
      "author_name": "humoyun",
      "author_url": "",
      "post_date": "05/30/2016 08:30:02",
      "content": "<pre><code>import csv\nimport json\n\nfilepath='D:/DATAFOLDER/kaggle_avito/ItemInfo_train.csv'\ncount=0\n\nwith open(filepath) as csvfile:\n    readCSV = csv.reader(csvfile, delimiter=',')\n    for row in readCSV:\n        # tmp = list(row)\n        # fmt=u'{:&lt;15}'*len(tmp)\n        # print fmt.format(*[s.decode('utf-8') for s in tmp])\n        print(&quot;[{0}] (0 row)  itemID       : {1} &quot;.format(count, row[0]))\n        print(&quot;[{0}] (1 row)  categoryID   : {1} &quot;.format(count, row[1]))\n        print(&quot;[{0}] (2 row)  title        : {1} &quot;.format(count, unicode(row[2])))\n        print(&quot;[{0}] (3 row)  description  : {1} &quot;.format(count, unicode(row[3])))\n        print(&quot;[{0}] (4 row)  images_array : {1} &quot;.format(count, row[4]))\n        print(&quot;[{0}] (5 row)  attrsJSON    : {1} &quot;.format(count, unicode(row[5])))\n        #data = json.loads(row[5])\n        print(&quot;[{0}] (6 row)  price        : {1} &quot;.format(count, row[6]))\n        print(&quot;[{0}] (7 row)  locationID   : {1} &quot;.format(count, row[7]))\n        print(&quot;[{0}] (8 row)  metroID      : {1} &quot;.format(count, row[8]))\n        print(&quot;[{0}] (9 row)  lat          : {1} &quot;.format(count, row[9]))\n        print(&quot;[{0}] (10 row) long         : {1} &quot;.format(count, row[10]))\n\n        print &quot;**********************************************************************************&quot;\n        count+=1\n</code></pre>\n\n<p>I was trying to use this script to explore data set, but It gives me <strong>UnicodeDecodeError</strong>: ascii codec can't decode byte 0xd0 in position...</p>\n\n<p>@ololo  Can you tweak my script so that I can read the file properly</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121833,
      "author_name": "agrigorev",
      "author_url": "",
      "post_date": "05/30/2016 08:37:51",
      "content": "<p>use codecs.open instead </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121836,
      "author_name": "humoyun",
      "author_url": "",
      "post_date": "05/30/2016 08:46:01",
      "content": "<p>@ololo same output</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "121811": "I am having difficulty of reading some fields of this file, namely, description, title and attrsJSON. Is there anybody who can provide some script on reading this fields properly?",
    "121821": "In python:\r\n\r\n    import pandas as pd\r\n    info = pd.read_csv( \"../input/ItemInfo_train.csv\", encoding=\"utf-8\")\r\n\r\nFor R, it may depend on the package you're using. I see solutions where they either pass it as a parameter (encoding='utf-8') or they change the locale for the script:\r\n\r\n    Sys.setlocale(category=\"LC_ALL\", locale = \"ru_RU\")\r\n\r\nor:\r\n\r\n    df_ch <- read.table(\"../input/ItemInfo_train.csv\",\r\n                         sep=\",\",\r\n                         header=FALSE, \r\n                         encoding=\"russian\", \r\n                         stringsAsFactors=FALSE\r\n                        )\r\n\r\nSomeone better experienced with R can probably give you a hint on which package should be used, they usually have different performances.",
    "121828": "import pandas as pd\r\n    info = pd.read_csv( \"../input/ItemInfo_train.csv\", encoding=\"utf-8\")\r\n\r\nBut this is taking too much time to load into memory, almost one minute, why is it extremely slow?",
    "121829": "humoyun, have a look at the file size",
    "121830": "import csv\r\n    import json\r\n    \r\n    filepath='D:/DATAFOLDER/kaggle_avito/ItemInfo_train.csv'\r\n    count=0\r\n    \r\n    with open(filepath) as csvfile:\r\n        readCSV = csv.reader(csvfile, delimiter=',')\r\n        for row in readCSV:\r\n        \t# tmp = list(row)\r\n        \t# fmt=u'{:<15}'*len(tmp)\r\n        \t# print fmt.format(*[s.decode('utf-8') for s in tmp])\r\n            print(\"[{0}] (0 row)  itemID \t   : {1} \".format(count, row[0]))\r\n            print(\"[{0}] (1 row)  categoryID   : {1} \".format(count, row[1]))\r\n            print(\"[{0}] (2 row)  title \t   : {1} \".format(count, unicode(row[2])))\r\n            print(\"[{0}] (3 row)  description  : {1} \".format(count, unicode(row[3])))\r\n            print(\"[{0}] (4 row)  images_array : {1} \".format(count, row[4]))\r\n            print(\"[{0}] (5 row)  attrsJSON    : {1} \".format(count, unicode(row[5])))\r\n            #data = json.loads(row[5])\r\n            print(\"[{0}] (6 row)  price        : {1} \".format(count, row[6]))\r\n            print(\"[{0}] (7 row)  locationID   : {1} \".format(count, row[7]))\r\n            print(\"[{0}] (8 row)  metroID  \t   : {1} \".format(count, row[8]))\r\n            print(\"[{0}] (9 row)  lat  \t\t   : {1} \".format(count, row[9]))\r\n            print(\"[{0}] (10 row) long         : {1} \".format(count, row[10]))\r\n         \r\n            print \"**********************************************************************************\"\r\n            count+=1\r\n\r\nI was trying to use this script to explore data set, but It gives me **UnicodeDecodeError**: ascii codec can't decode byte 0xd0 in position...\r\n\r\n@ololo  Can you tweak my script so that I can read the file properly\r\n\r\n\r\n\r\n\r\n  [1]: http://",
    "121833": "use codecs.open instead",
    "121836": "ololo same output"
  },
  "source": "meta"
}