{
  "id": 55511,
  "title": "Russian language not being recognized in CSV and also by R! ",
  "url": "/competitions/avito-demand-prediction/discussion/55511",
  "author_name": "",
  "post_date": "2018-04-27T12:18:21.734775900Z",
  "votes": null,
  "comment_count": 12,
  "views": 0,
  "content": "<p>I downloaded the train.csv file and the russian wordings are not clear when I open the CSV. Even the translator in R program interpreted it as Portuguese. Did anyone face this issue?</p>\n\n<p>For example: Region on first row in CSV is visible as 'Ð¡Ð²ÐµÑ€Ð´Ð»Ð¾Ð²ÑÐºÐ°Ñ Ð¾Ð±Ð»Ð°ÑÑ‚ÑŒ' in the actual csv file while the same on one of the Kernels shows 'Свердловская область'... why is excel changing the value? \nWhat am I doing wrong? :(</p>",
  "messages": [
    {
      "id": "320075",
      "postDate": "04/27/2018 12:18:21",
      "content": "<p>I downloaded the train.csv file and the russian wordings are not clear when I open the CSV. Even the translator in R program interpreted it as Portuguese. Did anyone face this issue?</p>\n\n<p>For example: Region on first row in CSV is visible as 'Ð¡Ð²ÐµÑ€Ð´Ð»Ð¾Ð²ÑÐºÐ°Ñ Ð¾Ð±Ð»Ð°ÑÑ‚ÑŒ' in the actual csv file while the same on one of the Kernels shows 'Свердловская область'... why is excel changing the value? \nWhat am I doing wrong? :(</p>",
      "rawMarkdown": "I downloaded the train.csv file and the russian wordings are not clear when I open the CSV. Even the translator in R program interpreted it as Portuguese. Did anyone face this issue?\n\nFor example: Region on first row in CSV is visible as 'Ð¡Ð²ÐµÑ€Ð´Ð»Ð¾Ð²ÑÐºÐ°Ñ Ð¾Ð±Ð»Ð°ÑÑ‚ÑŒ' in the actual csv file while the same on one of the Kernels shows 'Свердловская область'... why is excel changing the value? \nWhat am I doing wrong? :(",
      "votes": null
    },
    {
      "id": "320322",
      "postDate": "04/28/2018 09:45:52",
      "content": "<p>Not sure, but I had a similar problem - which was the reason I switched to Python for this one...</p>",
      "rawMarkdown": "Not sure, but I had a similar problem - which was the reason I switched to Python for this one...",
      "votes": null
    },
    {
      "id": "320357",
      "postDate": "04/28/2018 12:26:10",
      "content": "<p>But I get the same problem even when I open the csv file in excel. Did you face the same issue? \nFor example: Region on first row is visible as 'Ð¡Ð²ÐµÑ€Ð´Ð»Ð¾Ð²ÑÐºÐ°Ñ Ð¾Ð±Ð»Ð°ÑÑ‚ÑŒ' in the actual csv file while the same on one of the Kernels by SRK shows 'Свердловская область'... why is excel changing the value? Am I doing anything wrong? </p>",
      "rawMarkdown": "But I get the same problem even when I open the csv file in excel. Did you face the same issue? \nFor example: Region on first row is visible as 'Ð¡Ð²ÐµÑ€Ð´Ð»Ð¾Ð²ÑÐºÐ°Ñ Ð¾Ð±Ð»Ð°ÑÑ‚ÑŒ' in the actual csv file while the same on one of the Kernels by SRK shows 'Свердловская область'... why is excel changing the value? Am I doing anything wrong?",
      "votes": null
    },
    {
      "id": "320358",
      "postDate": "04/28/2018 12:31:47",
      "content": "<blockquote>\n  <p>when I open the CSV</p>\n</blockquote>\n\n<p>Unicode can be tricky to understand. Depending on the OS and many other things, the default encoding while opening the csv might not be the correct one for you. What you can try is to explicitly set the encoding, for example in Python try doing this: </p>\n\n<pre><code>train_data = pd.read_csv(\"../input/train.csv\", encoding=\"utf-8\")\n</code></pre>\n\n<p>Same goes for R.</p>",
      "rawMarkdown": "&gt; when I open the CSV\n\nUnicode can be tricky to understand. Depending on the OS and many other things, the default encoding while opening the csv might not be the correct one for you. What you can try is to explicitly set the encoding, for example in Python try doing this: \n\n    train_data = pd.read_csv(\"../input/train.csv\", encoding=\"utf-8\")\n\nSame goes for R.",
      "votes": null
    },
    {
      "id": "320554",
      "postDate": "04/29/2018 02:57:55",
      "content": "<p>Using encoding=\"utf-8\" didnt change anything in R... when I used encoding=\"UTF-8\", the columns were read as something like this'' though it didnt solve the problem... thanks for suggestion anyway, but do let me know if you sense there is something else I should try</p>",
      "rawMarkdown": "Using encoding=\"utf-8\" didnt change anything in R... when I used encoding=\"UTF-8\", the columns were read as something like this'",
      "votes": null
    },
    {
      "id": "320555",
      "postDate": "04/29/2018 03:00:55",
      "content": "<p>data=read.csv(\"C:/Users/VarunK/Desktop/Text_R/Kaggle/AVITO/train_active/train.csv\",nrows = 10,encoding=\"utf-8\"). \nHere is the code I used in R.</p>",
      "rawMarkdown": "data=read.csv(\"C:/Users/VarunK/Desktop/Text_R/Kaggle/AVITO/train_active/train.csv\",nrows = 10,encoding=\"utf-8\"). \nHere is the code I used in R.",
      "votes": null
    },
    {
      "id": "320583",
      "postDate": "04/29/2018 06:10:48",
      "content": "<p>Personally I do not use R, but I checked it out. I'd suggest you to try the <code>read_csv</code> function from the <code>readr</code> pacakage. First you can let it try to guess the encoding and see what it does:</p>\n\n<pre><code>data = read_csv(\"../input/train.csv\", n_max = 10)\n</code></pre>\n\n<p>If that doesn't work, then set the encoding explicitly like this:</p>\n\n<pre><code>data = read_csv(\"../input/train.csv\", n_max = 10, locale = locale(encoding = \"UTF-8\"))\n</code></pre>\n\n<p>Hopefully that would work, let me know if you get any problems still. In any case though, this thing is most probably about the proper encoding only.</p>",
      "rawMarkdown": "Personally I do not use R, but I checked it out. I'd suggest you to try the `read_csv` function from the `readr` pacakage. First you can let it try to guess the encoding and see what it does:\n\n    data = read_csv(\"../input/train.csv\", n_max = 10)\n\nIf that doesn't work, then set the encoding explicitly like this:\n\n    data = read_csv(\"../input/train.csv\", n_max = 10, locale = locale(encoding = \"UTF-8\"))\n\nHopefully that would work, let me know if you get any problems still. In any case though, this thing is most probably about the proper encoding only.",
      "votes": null
    },
    {
      "id": "320631",
      "postDate": "04/29/2018 10:47:44",
      "content": "<p>I had the same problem as you with other russian data. Spent a couple hours trying to import the data (at the time, SPSS file) using diff plugins, always with the same problem. What fixed it was changing the Locale to Russia/Russian and it worked.</p>",
      "rawMarkdown": "I had the same problem as you with other russian data. Spent a couple hours trying to import the data (at the time, SPSS file) using diff plugins, always with the same problem. What fixed it was changing the Locale to Russia/Russian and it worked.",
      "votes": null
    },
    {
      "id": "320636",
      "postDate": "04/29/2018 11:05:11",
      "content": "<p>Excel always does that. There's probably a way round somehow - but the file is too big for Excel anyway</p>",
      "rawMarkdown": "Excel always does that. There's probably a way round somehow - but the file is too big for Excel anyway",
      "votes": null
    },
    {
      "id": "320661",
      "postDate": "04/29/2018 12:32:19",
      "content": "<p>Wohoooo... this worked... the first comment worked :) Thanks Kishan... I hope to make some progress today...</p>",
      "rawMarkdown": "Wohoooo... this worked... the first comment worked :) Thanks Kishan... I hope to make some progress today...",
      "votes": null
    },
    {
      "id": "320662",
      "postDate": "04/29/2018 12:32:46",
      "content": "<p>Thanks Hugo... using read_csv worked</p>",
      "rawMarkdown": "Thanks Hugo... using read_csv worked",
      "votes": null
    },
    {
      "id": "321033",
      "postDate": "04/30/2018 12:45:23",
      "content": "<p>The <strong>read_csv()</strong> function from the <strong>tidyverse</strong> (or <strong>readr</strong>) package works for me.</p>\n\n<pre><code>library(tidyverse)\ntr &lt;- read_csv(\"../input/train.csv\")\n</code></pre>",
      "rawMarkdown": "The **read_csv()** function from the **tidyverse** (or **readr**) package works for me.\n\n    library(tidyverse)\n    tr &lt;- read_csv(\"../input/train.csv\")",
      "votes": null
    },
    {
      "id": "322197",
      "postDate": "05/02/2018 14:33:01",
      "content": "<p>We can change the location of language by</p>\n\n<pre><code>Sys.setlocale(,\"ru_RU\")\n</code></pre>\n\n<p>Further details can be found in <a href=\"https://stackoverflow.com/questions/14691555/cyrillic-encoding-output-in-r\">https://stackoverflow.com/questions/14691555/cyrillic-encoding-output-in-r</a></p>",
      "rawMarkdown": "We can change the location of language by\n\n    Sys.setlocale(,\"ru_RU\")\n\nFurther details can be found in https://stackoverflow.com/questions/14691555/cyrillic-encoding-output-in-r",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 320322,
      "author_name": "konradb",
      "author_url": "",
      "post_date": "04/28/2018 09:45:52",
      "content": "<p>Not sure, but I had a similar problem - which was the reason I switched to Python for this one...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 320357,
      "author_name": "vjkadekar",
      "author_url": "",
      "post_date": "04/28/2018 12:26:10",
      "content": "<p>But I get the same problem even when I open the csv file in excel. Did you face the same issue? \nFor example: Region on first row is visible as 'Ð¡Ð²ÐµÑ€Ð´Ð»Ð¾Ð²ÑÐºÐ°Ñ Ð¾Ð±Ð»Ð°ÑÑ‚ÑŒ' in the actual csv file while the same on one of the Kernels by SRK shows 'Свердловская область'... why is excel changing the value? Am I doing anything wrong? </p>",
      "votes": null,
      "replies": [
        {
          "id": 320636,
          "author_name": "domcastro",
          "author_url": "",
          "post_date": "04/29/2018 11:05:11",
          "content": "<p>Excel always does that. There's probably a way round somehow - but the file is too big for Excel anyway</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 320358,
      "author_name": "kishupro",
      "author_url": "",
      "post_date": "04/28/2018 12:31:47",
      "content": "<blockquote>\n  <p>when I open the CSV</p>\n</blockquote>\n\n<p>Unicode can be tricky to understand. Depending on the OS and many other things, the default encoding while opening the csv might not be the correct one for you. What you can try is to explicitly set the encoding, for example in Python try doing this: </p>\n\n<pre><code>train_data = pd.read_csv(\"../input/train.csv\", encoding=\"utf-8\")\n</code></pre>\n\n<p>Same goes for R.</p>",
      "votes": null,
      "replies": [
        {
          "id": 320554,
          "author_name": "vjkadekar",
          "author_url": "",
          "post_date": "04/29/2018 02:57:55",
          "content": "<p>Using encoding=\"utf-8\" didnt change anything in R... when I used encoding=\"UTF-8\", the columns were read as something like this'' though it didnt solve the problem... thanks for suggestion anyway, but do let me know if you sense there is something else I should try</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 320555,
          "author_name": "vjkadekar",
          "author_url": "",
          "post_date": "04/29/2018 03:00:55",
          "content": "<p>data=read.csv(\"C:/Users/VarunK/Desktop/Text_R/Kaggle/AVITO/train_active/train.csv\",nrows = 10,encoding=\"utf-8\"). \nHere is the code I used in R.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 320583,
          "author_name": "kishupro",
          "author_url": "",
          "post_date": "04/29/2018 06:10:48",
          "content": "<p>Personally I do not use R, but I checked it out. I'd suggest you to try the <code>read_csv</code> function from the <code>readr</code> pacakage. First you can let it try to guess the encoding and see what it does:</p>\n\n<pre><code>data = read_csv(\"../input/train.csv\", n_max = 10)\n</code></pre>\n\n<p>If that doesn't work, then set the encoding explicitly like this:</p>\n\n<pre><code>data = read_csv(\"../input/train.csv\", n_max = 10, locale = locale(encoding = \"UTF-8\"))\n</code></pre>\n\n<p>Hopefully that would work, let me know if you get any problems still. In any case though, this thing is most probably about the proper encoding only.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 320631,
          "author_name": "hugoncosta",
          "author_url": "",
          "post_date": "04/29/2018 10:47:44",
          "content": "<p>I had the same problem as you with other russian data. Spent a couple hours trying to import the data (at the time, SPSS file) using diff plugins, always with the same problem. What fixed it was changing the Locale to Russia/Russian and it worked.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 320661,
          "author_name": "vjkadekar",
          "author_url": "",
          "post_date": "04/29/2018 12:32:19",
          "content": "<p>Wohoooo... this worked... the first comment worked :) Thanks Kishan... I hope to make some progress today...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 320662,
          "author_name": "vjkadekar",
          "author_url": "",
          "post_date": "04/29/2018 12:32:46",
          "content": "<p>Thanks Hugo... using read_csv worked</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 321033,
      "author_name": "kailex",
      "author_url": "",
      "post_date": "04/30/2018 12:45:23",
      "content": "<p>The <strong>read_csv()</strong> function from the <strong>tidyverse</strong> (or <strong>readr</strong>) package works for me.</p>\n\n<pre><code>library(tidyverse)\ntr &lt;- read_csv(\"../input/train.csv\")\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 322197,
      "author_name": "sixianghu",
      "author_url": "",
      "post_date": "05/02/2018 14:33:01",
      "content": "<p>We can change the location of language by</p>\n\n<pre><code>Sys.setlocale(,\"ru_RU\")\n</code></pre>\n\n<p>Further details can be found in <a href=\"https://stackoverflow.com/questions/14691555/cyrillic-encoding-output-in-r\">https://stackoverflow.com/questions/14691555/cyrillic-encoding-output-in-r</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "320075": "I downloaded the train.csv file and the russian wordings are not clear when I open the CSV. Even the translator in R program interpreted it as Portuguese. Did anyone face this issue?\n\nFor example: Region on first row in CSV is visible as 'Ð¡Ð²ÐµÑ€Ð´Ð»Ð¾Ð²ÑÐºÐ°Ñ Ð¾Ð±Ð»Ð°ÑÑ‚ÑŒ' in the actual csv file while the same on one of the Kernels shows 'Свердловская область'... why is excel changing the value? \nWhat am I doing wrong? :(",
    "320322": "Not sure, but I had a similar problem - which was the reason I switched to Python for this one...",
    "320357": "But I get the same problem even when I open the csv file in excel. Did you face the same issue? \nFor example: Region on first row is visible as 'Ð¡Ð²ÐµÑ€Ð´Ð»Ð¾Ð²ÑÐºÐ°Ñ Ð¾Ð±Ð»Ð°ÑÑ‚ÑŒ' in the actual csv file while the same on one of the Kernels by SRK shows 'Свердловская область'... why is excel changing the value? Am I doing anything wrong?",
    "320358": "&gt; when I open the CSV\n\nUnicode can be tricky to understand. Depending on the OS and many other things, the default encoding while opening the csv might not be the correct one for you. What you can try is to explicitly set the encoding, for example in Python try doing this: \n\n    train_data = pd.read_csv(\"../input/train.csv\", encoding=\"utf-8\")\n\nSame goes for R.",
    "320554": "Using encoding=\"utf-8\" didnt change anything in R... when I used encoding=\"UTF-8\", the columns were read as something like this'",
    "320555": "data=read.csv(\"C:/Users/VarunK/Desktop/Text_R/Kaggle/AVITO/train_active/train.csv\",nrows = 10,encoding=\"utf-8\"). \nHere is the code I used in R.",
    "320583": "Personally I do not use R, but I checked it out. I'd suggest you to try the `read_csv` function from the `readr` pacakage. First you can let it try to guess the encoding and see what it does:\n\n    data = read_csv(\"../input/train.csv\", n_max = 10)\n\nIf that doesn't work, then set the encoding explicitly like this:\n\n    data = read_csv(\"../input/train.csv\", n_max = 10, locale = locale(encoding = \"UTF-8\"))\n\nHopefully that would work, let me know if you get any problems still. In any case though, this thing is most probably about the proper encoding only.",
    "320631": "I had the same problem as you with other russian data. Spent a couple hours trying to import the data (at the time, SPSS file) using diff plugins, always with the same problem. What fixed it was changing the Locale to Russia/Russian and it worked.",
    "320636": "Excel always does that. There's probably a way round somehow - but the file is too big for Excel anyway",
    "320661": "Wohoooo... this worked... the first comment worked :) Thanks Kishan... I hope to make some progress today...",
    "320662": "Thanks Hugo... using read_csv worked",
    "321033": "The **read_csv()** function from the **tidyverse** (or **readr**) package works for me.\n\n    library(tidyverse)\n    tr &lt;- read_csv(\"../input/train.csv\")",
    "322197": "We can change the location of language by\n\n    Sys.setlocale(,\"ru_RU\")\n\nFurther details can be found in https://stackoverflow.com/questions/14691555/cyrillic-encoding-output-in-r"
  },
  "source": "meta"
}