{
  "id": 6183,
  "title": "Handling 703,000,000 URLs",
  "url": "/competitions/yandex-personalized-web-search-challenge/discussion/6183",
  "author_name": "",
  "post_date": "2013-10-31T21:46:36.140Z",
  "votes": null,
  "comment_count": 4,
  "views": 1979,
  "content": "<p>I decided to use MongoDB for this effort.&nbsp; But after an hour of running, I only have about 300,000 URLs in my database.&nbsp;</p>\n<p>Either I am doing something wrong or it will take 97 days to load my DB.&nbsp; To that I say,&nbsp; Nyet!</p>\n<p>Does anyone know what the how the performance for MongoDB compares to Sqlite3 or PostGres ?&nbsp;</p>\n<p>BTW I am using Pypy with pymongo, so all code is jitted and fast, It is clearly hard drive bound.</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>",
  "messages": [
    {
      "id": "32973",
      "postDate": "10/31/2013 21:46:36",
      "content": "<p>I decided to use MongoDB for this effort.&nbsp; But after an hour of running, I only have about 300,000 URLs in my database.&nbsp;</p>\n<p>Either I am doing something wrong or it will take 97 days to load my DB.&nbsp; To that I say,&nbsp; Nyet!</p>\n<p>Does anyone know what the how the performance for MongoDB compares to Sqlite3 or PostGres ?&nbsp;</p>\n<p>BTW I am using Pypy with pymongo, so all code is jitted and fast, It is clearly hard drive bound.</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "33177",
      "postDate": "11/04/2013 15:16:27",
      "content": "<p>Whatever the form of database you're using for this, loading all URL's will take too much time. Even if you manage to get it you'll have large query execution times.</p>\n<p>I would strongly suggest that you sample the training set first and use that subset for building a model and only run your model calibration on the full data set once you're pretty confident that it produces reasonable results. Of course this comes at the expense of the extra work of coming up with a good sampling method to maintain the structure of the original dataset in the reduced set. Hope that helps.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "33212",
      "postDate": "11/04/2013 21:14:47",
      "content": "<p>Hi Bary,</p>\n<p>Agreed.&nbsp; I was running a test to see if my approach would scale.&nbsp; Obviously it does not.&nbsp;</p>\n<p>Interesting point about sampling the input file.&nbsp;&nbsp; I was going to use a percentage (maybe 5 days worth) for faster iteration and working out ideas.&nbsp;&nbsp;</p>\n<p>What advantages do you see in a more clever sampling method over a simplistic approach like taking the first 5 days?</p>\n<p>Thanks !</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "33256",
      "postDate": "11/05/2013 11:12:40",
      "content": "<p>Sampling a percentage of days in the training period is one way to do it. What it provides you with is a data subset that is rich in user information since you'll probably be preserving a large number of users but that reduced dataset will not be as rich in pointing out temporal variations (short period of time relative to the original training set). On the other hand, you could decide to take a sample of users instead of days, say 5% of the unique users in the dataset with all their associated queries over the entire period of time. This gives you the advantage of having the full view over the time period provided but with a reduced number of users.</p>\n<p>So it really depends on what your model is attempting to explore. If you're looking at temporal or time-dependent characteristics of querying habits for users, then you might be better off with a longer period of time to work with, hence sampling a set of users. If you're more interested in detailed user preference analysis and how that relates to query relevance, then you might want to have a large enough set of users and sampling a short period of time with a large number of users might be a better way of doing this.</p>\n<p>One more thing to bear in mind if you decide to sample is to watch out for your model's robustness. You want to make sure that your model performs on the full dataset close to how it performs on the reduced one. This is important because you're inevitably losing information in the sampling process. One way you could do that is to repeat the sampling process several times (say producing 5 reduced datasets) and comparing the performance of your models between these different subsets. You want the spread between performance across the different data samples to be as small as possible to ensure robustness.</p>\n<p>Hope that helps</p>\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "33265",
      "postDate": "11/05/2013 13:40:54",
      "content": "<p>&nbsp;</p>\n<p>In general I think users are learning as they search and explore.&nbsp; However, I don't believe I can capture this aspect given the limitations of the data set and the length of competition.&nbsp;&nbsp; So, my bias is to look at this month period as a single snapshot in time.&nbsp;</p>\n<p>The bigger challenge is in developing an objective performance measure because you don't want to optimize for the public leaderboard score.&nbsp; It is much easier to cross-validate under a supervised learning problem, where you know the answer beforehand.&nbsp; On my first pass reading of the DCG measure, I was left with the impression that I&nbsp; was missing some of the inputs (probably because they used the term &quot;Ideal DCG&quot;).&nbsp;&nbsp; I'll dig deeper into DCG to see if all the information is at my disposal.&nbsp; </p>\n<p>Thanks Bary, this has been a helpful train of thought...</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 33177,
      "author_name": "mohamedabdelbary",
      "author_url": "",
      "post_date": "11/04/2013 15:16:27",
      "content": "<p>Whatever the form of database you're using for this, loading all URL's will take too much time. Even if you manage to get it you'll have large query execution times.</p>\n<p>I would strongly suggest that you sample the training set first and use that subset for building a model and only run your model calibration on the full data set once you're pretty confident that it produces reasonable results. Of course this comes at the expense of the extra work of coming up with a good sampling method to maintain the structure of the original dataset in the reduced set. Hope that helps.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 33212,
      "author_name": "geringer",
      "author_url": "",
      "post_date": "11/04/2013 21:14:47",
      "content": "<p>Hi Bary,</p>\n<p>Agreed.&nbsp; I was running a test to see if my approach would scale.&nbsp; Obviously it does not.&nbsp;</p>\n<p>Interesting point about sampling the input file.&nbsp;&nbsp; I was going to use a percentage (maybe 5 days worth) for faster iteration and working out ideas.&nbsp;&nbsp;</p>\n<p>What advantages do you see in a more clever sampling method over a simplistic approach like taking the first 5 days?</p>\n<p>Thanks !</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 33256,
      "author_name": "mohamedabdelbary",
      "author_url": "",
      "post_date": "11/05/2013 11:12:40",
      "content": "<p>Sampling a percentage of days in the training period is one way to do it. What it provides you with is a data subset that is rich in user information since you'll probably be preserving a large number of users but that reduced dataset will not be as rich in pointing out temporal variations (short period of time relative to the original training set). On the other hand, you could decide to take a sample of users instead of days, say 5% of the unique users in the dataset with all their associated queries over the entire period of time. This gives you the advantage of having the full view over the time period provided but with a reduced number of users.</p>\n<p>So it really depends on what your model is attempting to explore. If you're looking at temporal or time-dependent characteristics of querying habits for users, then you might be better off with a longer period of time to work with, hence sampling a set of users. If you're more interested in detailed user preference analysis and how that relates to query relevance, then you might want to have a large enough set of users and sampling a short period of time with a large number of users might be a better way of doing this.</p>\n<p>One more thing to bear in mind if you decide to sample is to watch out for your model's robustness. You want to make sure that your model performs on the full dataset close to how it performs on the reduced one. This is important because you're inevitably losing information in the sampling process. One way you could do that is to repeat the sampling process several times (say producing 5 reduced datasets) and comparing the performance of your models between these different subsets. You want the spread between performance across the different data samples to be as small as possible to ensure robustness.</p>\n<p>Hope that helps</p>\n<p>&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 33265,
      "author_name": "geringer",
      "author_url": "",
      "post_date": "11/05/2013 13:40:54",
      "content": "<p>&nbsp;</p>\n<p>In general I think users are learning as they search and explore.&nbsp; However, I don't believe I can capture this aspect given the limitations of the data set and the length of competition.&nbsp;&nbsp; So, my bias is to look at this month period as a single snapshot in time.&nbsp;</p>\n<p>The bigger challenge is in developing an objective performance measure because you don't want to optimize for the public leaderboard score.&nbsp; It is much easier to cross-validate under a supervised learning problem, where you know the answer beforehand.&nbsp; On my first pass reading of the DCG measure, I was left with the impression that I&nbsp; was missing some of the inputs (probably because they used the term &quot;Ideal DCG&quot;).&nbsp;&nbsp; I'll dig deeper into DCG to see if all the information is at my disposal.&nbsp; </p>\n<p>Thanks Bary, this has been a helpful train of thought...</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "32973": "",
    "33177": "",
    "33212": "",
    "33256": "",
    "33265": ""
  },
  "source": "meta"
}