{
  "id": 15103,
  "title": "Is there easy ways to merge different files into one big file for classification",
  "url": "/competitions/avito-context-ad-clicks/discussion/15103",
  "author_name": "SkyLibrary",
  "post_date": "2015-07-07T23:06:09.583000",
  "votes": 1,
  "comment_count": 15,
  "views": 4172,
  "content": "<p>The problem is certainly interesting, but the dataset is too big, I can't find a way to use python or R to merge the datasets to do classification, it would be a much popular competition if there is some guidelines to prepare the data. </p>",
  "messages": [
    {
      "id": 84257,
      "postDate": "2015-07-13T07:00:13.330Z",
      "content": "<p>I merged them using sqlite. Here is my script:\n<a href=\"https://www.kaggle.com/mas313/avito-context-ad-clicks/merge-all-tables-into-one\">https://www.kaggle.com/mas313/avito-context-ad-clicks/merge-all-tables-into-one</a></p>\n\n<p>If you are a mac user then, sqlite3 is already installed on your machine. You can access it from terminal window by typing: sqlite3 PATHNAME/database.sql</p>\n\n<p>then you can copy paste the script into it. the output would be written in a cvs file: newdataset.csv  </p>",
      "rawMarkdown": "I merged them using sqlite. Here is my script:\r\nhttps://www.kaggle.com/mas313/avito-context-ad-clicks/merge-all-tables-into-one\r\n\r\nIf you are a mac user then, sqlite3 is already installed on your machine. You can access it from terminal window by typing: sqlite3 PATHNAME/database.sql\r\n\r\nthen you can copy paste the script into it. the output would be written in a cvs file: newdataset.csv  \r\n",
      "votes": 5
    },
    {
      "id": 83743,
      "postDate": "2015-07-08T02:18:52.637Z",
      "content": "<p>Just use sqlite is enough. Remember to create UNIQUE INDEX for IDs</p>\n\n<p>I'll post my scripts for combining these datasets. It's much faster than my previous ICDM scripts!</p>",
      "rawMarkdown": "Just use sqlite is enough. Remember to create UNIQUE INDEX for IDs\r\n\r\nI'll post my scripts for combining these datasets. It's much faster than my previous ICDM scripts!",
      "votes": 5
    },
    {
      "id": 86792,
      "postDate": "2015-07-24T11:53:48.173Z",
      "content": "<p>If you are on linux, try sorting the files by the common identifier next. With that you you have to load fewer records into memory.</p>\n\n<p>[quote=Jesse Burstr&#246;m;86781]</p>\n\n<p>I tried R and Python to merge SearchStream with SearchInfo but just could not get any version of 'line by line' or 'block by block' implementation to work without running out of memory. I tried awk to get the SearchID from SearchInfo into one file (1GB) but just could not load it into R or Python due to some entry beeing non-numeric (due to corrupt line) -&gt; causes memory out failure. I think the scripting languages tends to defragment the memory and/or creates a history heap that inflates during repeated processing. </p>\n\n<p>Since all good solutions seems to use the sqlite database i try that instead. Interesting to learn all sorts about large csv files and how they process after all but a bit frustrating since not much statistical analysis has been made, but i will give a last go at ftrl using sqlite...</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "If you are on linux, try sorting the files by the common identifier next. With that you you have to load fewer records into memory.\r\n\r\n[quote=Jesse Burström;86781]\r\n\r\nI tried R and Python to merge SearchStream with SearchInfo but just could not get any version of 'line by line' or 'block by block' implementation to work without running out of memory. I tried awk to get the SearchID from SearchInfo into one file (1GB) but just could not load it into R or Python due to some entry beeing non-numeric (due to corrupt line) -> causes memory out failure. I think the scripting languages tends to defragment the memory and/or creates a history heap that inflates during repeated processing. \r\n\r\nSince all good solutions seems to use the sqlite database i try that instead. Interesting to learn all sorts about large csv files and how they process after all but a bit frustrating since not much statistical analysis has been made, but i will give a last go at ftrl using sqlite...\r\n\r\n[/quote]\r\n",
      "votes": 1
    },
    {
      "id": 83934,
      "postDate": "2015-07-09T13:30:45.320Z",
      "content": "<p>I decide not to post my whole script. Since it turns out give me top 20 LB score using ftrl without tuning, which I'm afraid too good to make public at this point(only 19 days left).</p>\n\n<p>BTW, I use the sqlite file, and create UNIQUE INDEX for SearchInfo, UserInfo, AdsInfo, Category and Location. Then combine them into SearchStream. I use pypy to accelerate it, and it costs about 3 hours to run.</p>",
      "rawMarkdown": "I decide not to post my whole script. Since it turns out give me top 20 LB score using ftrl without tuning, which I'm afraid too good to make public at this point(only 19 days left).\r\n\r\nBTW, I use the sqlite file, and create UNIQUE INDEX for SearchInfo, UserInfo, AdsInfo, Category and Location. Then combine them into SearchStream. I use pypy to accelerate it, and it costs about 3 hours to run.",
      "votes": 1
    },
    {
      "id": 83732,
      "postDate": "2015-07-08T00:25:30.650Z",
      "content": "<p>I have been reading the file line-by-line so far - it is a linear time process so works OK with any memory constrains, but it definitely is not optimal.  I will try SQL next, and I am sure that will handle data of this size competently.</p>",
      "rawMarkdown": "I have been reading the file line-by-line so far - it is a linear time process so works OK with any memory constrains, but it definitely is not optimal.  I will try SQL next, and I am sure that will handle data of this size competently.",
      "votes": 1
    },
    {
      "id": 83721,
      "postDate": "2015-07-07T23:06:09.583Z",
      "content": "<p>The problem is certainly interesting, but the dataset is too big, I can't find a way to use python or R to merge the datasets to do classification, it would be a much popular competition if there is some guidelines to prepare the data. </p>",
      "rawMarkdown": "The problem is certainly interesting, but the dataset is too big, I can't find a way to use python or R to merge the datasets to do classification, it would be a much popular competition if there is some guidelines to prepare the data. ",
      "votes": 1
    },
    {
      "id": 83895,
      "postDate": "2015-07-09T06:09:40.787Z",
      "content": "<p>The two big files are both sorted by the same ID, the easiest way to merge data without the need to have more than two lines of text in memory at the same time is:\nRead lines from trainSearchStream.tsv\nwhen ever there is a new SearchID just read lines from SearchInfo.tsv untill you get to that ID</p>\n\n<p>Smaller tables can easily be put into memory and looked up when needed.  Or you can make some simple database to do lookups on the other tables if your very low and can't even fit the small tables into memory.</p>\n\n<p>Then again I don't know how memory efficient R and Python are. I write all my code in java and can just choose the optimal structures to minimize memory usage.</p>",
      "rawMarkdown": "The two big files are both sorted by the same ID, the easiest way to merge data without the need to have more than two lines of text in memory at the same time is:\r\nRead lines from trainSearchStream.tsv\r\nwhen ever there is a new SearchID just read lines from SearchInfo.tsv untill you get to that ID\r\n\r\nSmaller tables can easily be put into memory and looked up when needed.  Or you can make some simple database to do lookups on the other tables if your very low and can't even fit the small tables into memory.\r\n\r\nThen again I don't know how memory efficient R and Python are. I write all my code in java and can just choose the optimal structures to minimize memory usage.",
      "votes": 2
    },
    {
      "id": 83882,
      "postDate": "2015-07-09T01:53:09.420Z",
      "content": "<p>Very small, less than 500 MB. I'm still checking the output. The model built upon my combined data is worse than benchmarks. This is abnormal since my data has much more meaningful features, so I guess there might be some mistakes.</p>",
      "rawMarkdown": "Very small, less than 500 MB. I'm still checking the output. The model built upon my combined data is worse than benchmarks. This is abnormal since my data has much more meaningful features, so I guess there might be some mistakes.",
      "votes": 2
    },
    {
      "id": 86199,
      "postDate": "2015-07-20T21:14:47.557Z",
      "content": "<p>[quote=Stephen McInerney;83972]</p>\n\n<p>R and Python are as memory-efficient as any other language; we're talking about SQL join operations.\n(Python transaction-handlers are useful.)</p>\n\n<p>Specifically, the gating factor here would be <strong>multijoins across multiple huge tables</strong>. Any algorithm which avoids that will save lots of memory.</p>\n\n<p>Anyway, Jiming's 3hrs runtime for a suboptimal baseline sounds like too much pain for me so I'll skip this competition. Good luck.</p>\n\n<p>[/quote]</p>\n\n<p>Here is good information on why one should use Pandas instead of sqlite3 or data.table (assuming enough memory is available) :</p>\n\n<p><a href=\"http://wesmckinney.com/blog/high-performance-database-joins-with-pandas-dataframe-more-benchmarks/\">http://wesmckinney.com/blog/high-performance-database-joins-with-pandas-dataframe-more-benchmarks/</a></p>\n\n<p>That being said, I bet Spark joins would be even faster.</p>",
      "rawMarkdown": "[quote=Stephen McInerney;83972]\r\n\r\nR and Python are as memory-efficient as any other language; we're talking about SQL join operations.\r\n(Python transaction-handlers are useful.)\r\n\r\nSpecifically, the gating factor here would be **multijoins across multiple huge tables**. Any algorithm which avoids that will save lots of memory.\r\n\r\nAnyway, Jiming's 3hrs runtime for a suboptimal baseline sounds like too much pain for me so I'll skip this competition. Good luck.\r\n\r\n[/quote]\r\n\r\nHere is good information on why one should use Pandas instead of sqlite3 or data.table (assuming enough memory is available) :\r\n\r\nhttp://wesmckinney.com/blog/high-performance-database-joins-with-pandas-dataframe-more-benchmarks/\r\n\r\nThat being said, I bet Spark joins would be even faster.\r\n"
    },
    {
      "id": 85998,
      "postDate": "2015-07-19T01:21:12.010Z",
      "content": "<p>I believe the sqlite solution is the best one for this competition since there are already guidelines on how to implement it. Personally, I prefer postgreSQL. <a href=\"http://www.postgresql.org/download/\">http://www.postgresql.org/download/</a></p>",
      "rawMarkdown": "I believe the sqlite solution is the best one for this competition since there are already guidelines on how to implement it. Personally, I prefer postgreSQL. http://www.postgresql.org/download/"
    },
    {
      "id": 83972,
      "postDate": "2015-07-09T20:38:04.740Z",
      "content": "<p>R and Python are as memory-efficient as any other language; we're talking about SQL join operations.\n(Python transaction-handlers are useful.)</p>\n\n<p>Specifically, the gating factor here would be <strong>multijoins across multiple huge tables</strong>. Any algorithm which avoids that will save lots of memory.</p>\n\n<p>Anyway, Jiming's 3hrs runtime for a suboptimal baseline sounds like too much pain for me so I'll skip this competition. Good luck.</p>",
      "rawMarkdown": "R and Python are as memory-efficient as any other language; we're talking about SQL join operations.\r\n(Python transaction-handlers are useful.)\r\n\r\nSpecifically, the gating factor here would be **multijoins across multiple huge tables**. Any algorithm which avoids that will save lots of memory.\r\n\r\nAnyway, Jiming's 3hrs runtime for a suboptimal baseline sounds like too much pain for me so I'll skip this competition. Good luck."
    },
    {
      "id": 83878,
      "postDate": "2015-07-09T01:20:18.713Z",
      "content": "<p>Jiming et al - how much memory footprint do you need to process the datasets? 32 Gb? more?</p>",
      "rawMarkdown": "Jiming et al - how much memory footprint do you need to process the datasets? 32 Gb? more?"
    },
    {
      "id": 83757,
      "postDate": "2015-07-08T04:49:13.983Z",
      "content": "<p>30G is not enough. ICDM is also interesting, I'll turn to there once I finish this. </p>",
      "rawMarkdown": "30G is not enough. ICDM is also interesting, I'll turn to there once I finish this. "
    },
    {
      "id": 83756,
      "postDate": "2015-07-08T04:31:00.287Z",
      "content": "<p>I might not enter this competition with only 20 days left, the dataset is too big for my laptop, I just checked, my laptop now only have 30G free space left, I am not even sure if I can unzip these files or not, I might go to ICDM instead, which still have 40+ days left. \nBut thanks guys for the guide, I still like to learn something new from this competition.</p>",
      "rawMarkdown": "I might not enter this competition with only 20 days left, the dataset is too big for my laptop, I just checked, my laptop now only have 30G free space left, I am not even sure if I can unzip these files or not, I might go to ICDM instead, which still have 40+ days left. \r\nBut thanks guys for the guide, I still like to learn something new from this competition."
    },
    {
      "id": 86863,
      "postDate": "2015-07-24T18:46:23.043Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 86781,
      "postDate": "2015-07-24T09:58:05.910Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 84257,
      "author_name": "MAS",
      "author_url": "",
      "post_date": "2015-07-13T07:00:13.330000",
      "content": "<p>I merged them using sqlite. Here is my script:\n<a href=\"https://www.kaggle.com/mas313/avito-context-ad-clicks/merge-all-tables-into-one\">https://www.kaggle.com/mas313/avito-context-ad-clicks/merge-all-tables-into-one</a></p>\n\n<p>If you are a mac user then, sqlite3 is already installed on your machine. You can access it from terminal window by typing: sqlite3 PATHNAME/database.sql</p>\n\n<p>then you can copy paste the script into it. the output would be written in a cvs file: newdataset.csv  </p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 83743,
      "author_name": "Jiming Ye",
      "author_url": "",
      "post_date": "2015-07-08T02:18:52.637000",
      "content": "<p>Just use sqlite is enough. Remember to create UNIQUE INDEX for IDs</p>\n\n<p>I'll post my scripts for combining these datasets. It's much faster than my previous ICDM scripts!</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 86792,
      "author_name": "Leustagos",
      "author_url": "",
      "post_date": "2015-07-24T11:53:48.173000",
      "content": "<p>If you are on linux, try sorting the files by the common identifier next. With that you you have to load fewer records into memory.</p>\n\n<p>[quote=Jesse Burstr&#246;m;86781]</p>\n\n<p>I tried R and Python to merge SearchStream with SearchInfo but just could not get any version of 'line by line' or 'block by block' implementation to work without running out of memory. I tried awk to get the SearchID from SearchInfo into one file (1GB) but just could not load it into R or Python due to some entry beeing non-numeric (due to corrupt line) -&gt; causes memory out failure. I think the scripting languages tends to defragment the memory and/or creates a history heap that inflates during repeated processing. </p>\n\n<p>Since all good solutions seems to use the sqlite database i try that instead. Interesting to learn all sorts about large csv files and how they process after all but a bit frustrating since not much statistical analysis has been made, but i will give a last go at ftrl using sqlite...</p>\n\n<p>[/quote]</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 83934,
      "author_name": "Jiming Ye",
      "author_url": "",
      "post_date": "2015-07-09T13:30:45.320000",
      "content": "<p>I decide not to post my whole script. Since it turns out give me top 20 LB score using ftrl without tuning, which I'm afraid too good to make public at this point(only 19 days left).</p>\n\n<p>BTW, I use the sqlite file, and create UNIQUE INDEX for SearchInfo, UserInfo, AdsInfo, Category and Location. Then combine them into SearchStream. I use pypy to accelerate it, and it costs about 3 hours to run.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 83732,
      "author_name": "DerekZH",
      "author_url": "",
      "post_date": "2015-07-08T00:25:30.650000",
      "content": "<p>I have been reading the file line-by-line so far - it is a linear time process so works OK with any memory constrains, but it definitely is not optimal.  I will try SQL next, and I am sure that will handle data of this size competently.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 83895,
      "author_name": "Lars Ropeid Selsås",
      "author_url": "",
      "post_date": "2015-07-09T06:09:40.787000",
      "content": "<p>The two big files are both sorted by the same ID, the easiest way to merge data without the need to have more than two lines of text in memory at the same time is:\nRead lines from trainSearchStream.tsv\nwhen ever there is a new SearchID just read lines from SearchInfo.tsv untill you get to that ID</p>\n\n<p>Smaller tables can easily be put into memory and looked up when needed.  Or you can make some simple database to do lookups on the other tables if your very low and can't even fit the small tables into memory.</p>\n\n<p>Then again I don't know how memory efficient R and Python are. I write all my code in java and can just choose the optimal structures to minimize memory usage.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 83882,
      "author_name": "Jiming Ye",
      "author_url": "",
      "post_date": "2015-07-09T01:53:09.420000",
      "content": "<p>Very small, less than 500 MB. I'm still checking the output. The model built upon my combined data is worse than benchmarks. This is abnormal since my data has much more meaningful features, so I guess there might be some mistakes.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 86199,
      "author_name": "quackTau",
      "author_url": "",
      "post_date": "2015-07-20T21:14:47.557000",
      "content": "<p>[quote=Stephen McInerney;83972]</p>\n\n<p>R and Python are as memory-efficient as any other language; we're talking about SQL join operations.\n(Python transaction-handlers are useful.)</p>\n\n<p>Specifically, the gating factor here would be <strong>multijoins across multiple huge tables</strong>. Any algorithm which avoids that will save lots of memory.</p>\n\n<p>Anyway, Jiming's 3hrs runtime for a suboptimal baseline sounds like too much pain for me so I'll skip this competition. Good luck.</p>\n\n<p>[/quote]</p>\n\n<p>Here is good information on why one should use Pandas instead of sqlite3 or data.table (assuming enough memory is available) :</p>\n\n<p><a href=\"http://wesmckinney.com/blog/high-performance-database-joins-with-pandas-dataframe-more-benchmarks/\">http://wesmckinney.com/blog/high-performance-database-joins-with-pandas-dataframe-more-benchmarks/</a></p>\n\n<p>That being said, I bet Spark joins would be even faster.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 85998,
      "author_name": "Jordan Sundheim",
      "author_url": "",
      "post_date": "2015-07-19T01:21:12.010000",
      "content": "<p>I believe the sqlite solution is the best one for this competition since there are already guidelines on how to implement it. Personally, I prefer postgreSQL. <a href=\"http://www.postgresql.org/download/\">http://www.postgresql.org/download/</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 83972,
      "author_name": "Stephen McInerney",
      "author_url": "",
      "post_date": "2015-07-09T20:38:04.740000",
      "content": "<p>R and Python are as memory-efficient as any other language; we're talking about SQL join operations.\n(Python transaction-handlers are useful.)</p>\n\n<p>Specifically, the gating factor here would be <strong>multijoins across multiple huge tables</strong>. Any algorithm which avoids that will save lots of memory.</p>\n\n<p>Anyway, Jiming's 3hrs runtime for a suboptimal baseline sounds like too much pain for me so I'll skip this competition. Good luck.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 83878,
      "author_name": "Stephen McInerney",
      "author_url": "",
      "post_date": "2015-07-09T01:20:18.713000",
      "content": "<p>Jiming et al - how much memory footprint do you need to process the datasets? 32 Gb? more?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 83757,
      "author_name": "Jiming Ye",
      "author_url": "",
      "post_date": "2015-07-08T04:49:13.983000",
      "content": "<p>30G is not enough. ICDM is also interesting, I'll turn to there once I finish this. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 83756,
      "author_name": "SkyLibrary",
      "author_url": "",
      "post_date": "2015-07-08T04:31:00.287000",
      "content": "<p>I might not enter this competition with only 20 days left, the dataset is too big for my laptop, I just checked, my laptop now only have 30G free space left, I am not even sure if I can unzip these files or not, I might go to ICDM instead, which still have 40+ days left. \nBut thanks guys for the guide, I still like to learn something new from this competition.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 86863,
      "author_name": "",
      "author_url": "",
      "post_date": "2015-07-24T18:46:23.043000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 86781,
      "author_name": "",
      "author_url": "",
      "post_date": "2015-07-24T09:58:05.910000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "84257": "I merged them using sqlite. Here is my script:\r\nhttps://www.kaggle.com/mas313/avito-context-ad-clicks/merge-all-tables-into-one\r\n\r\nIf you are a mac user then, sqlite3 is already installed on your machine. You can access it from terminal window by typing: sqlite3 PATHNAME/database.sql\r\n\r\nthen you can copy paste the script into it. the output would be written in a cvs file: newdataset.csv  \r\n",
    "83743": "Just use sqlite is enough. Remember to create UNIQUE INDEX for IDs\r\n\r\nI'll post my scripts for combining these datasets. It's much faster than my previous ICDM scripts!",
    "86792": "If you are on linux, try sorting the files by the common identifier next. With that you you have to load fewer records into memory.\r\n\r\n[quote=Jesse Burström;86781]\r\n\r\nI tried R and Python to merge SearchStream with SearchInfo but just could not get any version of 'line by line' or 'block by block' implementation to work without running out of memory. I tried awk to get the SearchID from SearchInfo into one file (1GB) but just could not load it into R or Python due to some entry beeing non-numeric (due to corrupt line) -> causes memory out failure. I think the scripting languages tends to defragment the memory and/or creates a history heap that inflates during repeated processing. \r\n\r\nSince all good solutions seems to use the sqlite database i try that instead. Interesting to learn all sorts about large csv files and how they process after all but a bit frustrating since not much statistical analysis has been made, but i will give a last go at ftrl using sqlite...\r\n\r\n[/quote]\r\n",
    "83934": "I decide not to post my whole script. Since it turns out give me top 20 LB score using ftrl without tuning, which I'm afraid too good to make public at this point(only 19 days left).\r\n\r\nBTW, I use the sqlite file, and create UNIQUE INDEX for SearchInfo, UserInfo, AdsInfo, Category and Location. Then combine them into SearchStream. I use pypy to accelerate it, and it costs about 3 hours to run.",
    "83732": "I have been reading the file line-by-line so far - it is a linear time process so works OK with any memory constrains, but it definitely is not optimal.  I will try SQL next, and I am sure that will handle data of this size competently.",
    "83721": "The problem is certainly interesting, but the dataset is too big, I can't find a way to use python or R to merge the datasets to do classification, it would be a much popular competition if there is some guidelines to prepare the data. ",
    "83895": "The two big files are both sorted by the same ID, the easiest way to merge data without the need to have more than two lines of text in memory at the same time is:\r\nRead lines from trainSearchStream.tsv\r\nwhen ever there is a new SearchID just read lines from SearchInfo.tsv untill you get to that ID\r\n\r\nSmaller tables can easily be put into memory and looked up when needed.  Or you can make some simple database to do lookups on the other tables if your very low and can't even fit the small tables into memory.\r\n\r\nThen again I don't know how memory efficient R and Python are. I write all my code in java and can just choose the optimal structures to minimize memory usage.",
    "83882": "Very small, less than 500 MB. I'm still checking the output. The model built upon my combined data is worse than benchmarks. This is abnormal since my data has much more meaningful features, so I guess there might be some mistakes.",
    "86199": "[quote=Stephen McInerney;83972]\r\n\r\nR and Python are as memory-efficient as any other language; we're talking about SQL join operations.\r\n(Python transaction-handlers are useful.)\r\n\r\nSpecifically, the gating factor here would be **multijoins across multiple huge tables**. Any algorithm which avoids that will save lots of memory.\r\n\r\nAnyway, Jiming's 3hrs runtime for a suboptimal baseline sounds like too much pain for me so I'll skip this competition. Good luck.\r\n\r\n[/quote]\r\n\r\nHere is good information on why one should use Pandas instead of sqlite3 or data.table (assuming enough memory is available) :\r\n\r\nhttp://wesmckinney.com/blog/high-performance-database-joins-with-pandas-dataframe-more-benchmarks/\r\n\r\nThat being said, I bet Spark joins would be even faster.\r\n",
    "85998": "I believe the sqlite solution is the best one for this competition since there are already guidelines on how to implement it. Personally, I prefer postgreSQL. http://www.postgresql.org/download/",
    "83972": "R and Python are as memory-efficient as any other language; we're talking about SQL join operations.\r\n(Python transaction-handlers are useful.)\r\n\r\nSpecifically, the gating factor here would be **multijoins across multiple huge tables**. Any algorithm which avoids that will save lots of memory.\r\n\r\nAnyway, Jiming's 3hrs runtime for a suboptimal baseline sounds like too much pain for me so I'll skip this competition. Good luck.",
    "83878": "Jiming et al - how much memory footprint do you need to process the datasets? 32 Gb? more?",
    "83757": "30G is not enough. ICDM is also interesting, I'll turn to there once I finish this. ",
    "83756": "I might not enter this competition with only 20 days left, the dataset is too big for my laptop, I just checked, my laptop now only have 30G free space left, I am not even sure if I can unzip these files or not, I might go to ICDM instead, which still have 40+ days left. \r\nBut thanks guys for the guide, I still like to learn something new from this competition.",
    "86863": "",
    "86781": ""
  }
}