{
  "id": 14592,
  "title": "Postgres – Import Data – SQL Script",
  "url": "/competitions/avito-context-ad-clicks/discussion/14592",
  "author_name": "",
  "post_date": "2015-06-06T20:51:20.637Z",
  "votes": 9,
  "comment_count": 1,
  "views": 1087,
  "content": "<p>In case somebody has a lot of disk space ;) and wants&nbsp;to load the data to Postgres I attach a script that creates all the tables with data types optimized for what the columns contain (smallint or bit if posstible, or varchar of minimal size). Also, tables that have rpimary keys are clustered on them, and analyzed. Created indexes have <em>FILLFACTOR = 100</em> as I assume I will not update those tables. I do not remember how long does it take to load everything and create indexes as I run them step by step while writing the script, but I guess it will take several hours (I have an SSD drive).</p>\n<p>One more thing, if you want to use the script make sure you change paths to files. Also you may want to remove the <em>TABLESPACE</em> &#8211; I have two SSD drives, and to save space I have some tables on one of them and some on the other one<em>.&nbsp;</em>The whole dataset after importing, sorting, indexing takes about 79GB (indexes are quite heavy). Postgres as such does not allow for arbitrary compression. It decides by itself, another option is to enable it on partition &#8211; I have not.</p>\n<p>If you need any more information about the script/DB feel free to ask. Also if you have any tips how to improve the DB's querying perfromance I will be more than happy to learn something new. :)</p>\n<p><strong>NOTE</strong>: as I understand dictionaries that appear in some of the datasets cannot be stored in Postgres as JSON as the DB does not recognize the numerical keys properly &#8211; it expects them to be quoted strings. It makes things a bit harder to work with, but not impossible to query.</p>\n<p><strong>EDIT</strong>: I have no idea why there are two versions of the script attached, I added one line to the first one, removed it, and attached again. I expected to see only one of them, but it looks like both of them were uploaded and attached to the post.</p>\n\n<p><strong>EDIT2</strong>: Sorry for the mess, I missed one semicolon in the code and one table. New version, 0.2 attached.</p>",
  "messages": [
    {
      "id": "81137",
      "postDate": "06/06/2015 20:51:20",
      "content": "<p>In case somebody has a lot of disk space ;) and wants&nbsp;to load the data to Postgres I attach a script that creates all the tables with data types optimized for what the columns contain (smallint or bit if posstible, or varchar of minimal size). Also, tables that have rpimary keys are clustered on them, and analyzed. Created indexes have <em>FILLFACTOR = 100</em> as I assume I will not update those tables. I do not remember how long does it take to load everything and create indexes as I run them step by step while writing the script, but I guess it will take several hours (I have an SSD drive).</p>\n<p>One more thing, if you want to use the script make sure you change paths to files. Also you may want to remove the <em>TABLESPACE</em> &#8211; I have two SSD drives, and to save space I have some tables on one of them and some on the other one<em>.&nbsp;</em>The whole dataset after importing, sorting, indexing takes about 79GB (indexes are quite heavy). Postgres as such does not allow for arbitrary compression. It decides by itself, another option is to enable it on partition &#8211; I have not.</p>\n<p>If you need any more information about the script/DB feel free to ask. Also if you have any tips how to improve the DB's querying perfromance I will be more than happy to learn something new. :)</p>\n<p><strong>NOTE</strong>: as I understand dictionaries that appear in some of the datasets cannot be stored in Postgres as JSON as the DB does not recognize the numerical keys properly &#8211; it expects them to be quoted strings. It makes things a bit harder to work with, but not impossible to query.</p>\n<p><strong>EDIT</strong>: I have no idea why there are two versions of the script attached, I added one line to the first one, removed it, and attached again. I expected to see only one of them, but it looks like both of them were uploaded and attached to the post.</p>\n\n<p><strong>EDIT2</strong>: Sorry for the mess, I missed one semicolon in the code and one table. New version, 0.2 attached.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "82500",
      "postDate": "06/21/2015 12:45:46",
      "content": "<p>Thanks mpekalski! Very useful. For the ones who struggle to get the Russian language into R and find that the encoding might be a horror on a windows machine. This works for me:</p>\n<p><em>library(RPostgreSQL)</em></p>\n<p><em>channel &lt;- dbConnect(PostgreSQL(),user= &quot;xxxxx&quot;, password=&quot;xxxxx&quot;, dbname=&quot;xxxxx&quot;)</em></p>\n<p><em>data &lt;- dbGetQuery(channel,&quot;SELECT * FROM ads_info LIMIT 10&quot;)</em></p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 82500,
      "author_name": "mennotaanman",
      "author_url": "",
      "post_date": "06/21/2015 12:45:46",
      "content": "<p>Thanks mpekalski! Very useful. For the ones who struggle to get the Russian language into R and find that the encoding might be a horror on a windows machine. This works for me:</p>\n<p><em>library(RPostgreSQL)</em></p>\n<p><em>channel &lt;- dbConnect(PostgreSQL(),user= &quot;xxxxx&quot;, password=&quot;xxxxx&quot;, dbname=&quot;xxxxx&quot;)</em></p>\n<p><em>data &lt;- dbGetQuery(channel,&quot;SELECT * FROM ads_info LIMIT 10&quot;)</em></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "81137": "",
    "82500": ""
  },
  "source": "meta"
}