{
  "id": 24844,
  "title": "Thinks I've tried that are no better than BTB (0.635)",
  "url": "/competitions/outbrain-click-prediction/discussion/24844",
  "author_name": "",
  "post_date": "2016-10-27T21:34:48.010Z",
  "votes": 6,
  "comment_count": 2,
  "views": 669,
  "content": "<p>I've tried a number of different approach, none of which is better than the simple Popularity Recommender first implemented by Clustifier.</p>\n\n<p>Since this is a large data set I'm using <a href=\"https://turi.com/\">GraphLab</a> and AWS spot instance with up to 32 cores and 64GB memory. This has allowed me to try the following:</p>\n\n<p>1) ** <a href=\"https://turi.com/products/create/docs/generated/graphlab.recommender.item_similarity_recommender.ItemSimilarityRecommender.html\">Popularity Recommender</a>**\nThis is the same algorithm as used in BTB by Clustifier but optimised in the GL library. It's one line of code and takes less than 2 min to create a model. My CV score was <strong>0.6442</strong> and the submission was  0.6352 (i.e., identical to the BTB). </p>\n\n<p>2) ** <a href=\"https://turi.com/products/create/docs/generated/graphlab.recommender.item_similarity_recommender.ItemSimilarityRecommender.html\">Popularity Recommender</a>**\nThis ranks items according to its similarity to other items observed for the user in question.  For this I only used the training data where clicked ==1. This took less than a min to run and the CV score was <strong>0.5219</strong></p>\n\n<p>3) ** <a href=\"https://turi.com/products/create/docs/generated/graphlab.recommender.item_content_recommender.ItemContentRecommender.html\">Item Content Recommender</a> \nThis algorithm computes the similarity between items using the content of each item.\nThis uses  nearest neighbours and in many cases took too long. Once simple version  took about 30 min and had a CV Score of <strong>0.4617</strong> at which point I gave up.</p>\n\n<p>4) ** <a href=\"https://turi.com/products/create/docs/generated/graphlab.recommender.factorization_recommender.FactorizationRecommender.html\">Factorization Recommender</a>\nThis learns latent factors for each user and item and uses them to make predictions.</p>\n\n<p>This took about 10 min and the best CV was 0.5965</p>\n\n<p>5) Web page to Ad page similarity\nFor this I calculated the cosine similarity between the document metadata (i.e., topic, category, entity) for the web page the ads were shown in vs the web page the ad linked to, then used a boosted tree make predictions.  </p>\n\n<p>Since this required a lot of joins I used a 10% subset of the data and the best CV was 0.480</p>",
  "messages": [
    {
      "id": "141681",
      "postDate": "10/27/2016 21:34:48",
      "content": "<p>I've tried a number of different approach, none of which is better than the simple Popularity Recommender first implemented by Clustifier.</p>\n\n<p>Since this is a large data set I'm using <a href=\"https://turi.com/\">GraphLab</a> and AWS spot instance with up to 32 cores and 64GB memory. This has allowed me to try the following:</p>\n\n<p>1) ** <a href=\"https://turi.com/products/create/docs/generated/graphlab.recommender.item_similarity_recommender.ItemSimilarityRecommender.html\">Popularity Recommender</a>**\nThis is the same algorithm as used in BTB by Clustifier but optimised in the GL library. It's one line of code and takes less than 2 min to create a model. My CV score was <strong>0.6442</strong> and the submission was  0.6352 (i.e., identical to the BTB). </p>\n\n<p>2) ** <a href=\"https://turi.com/products/create/docs/generated/graphlab.recommender.item_similarity_recommender.ItemSimilarityRecommender.html\">Popularity Recommender</a>**\nThis ranks items according to its similarity to other items observed for the user in question.  For this I only used the training data where clicked ==1. This took less than a min to run and the CV score was <strong>0.5219</strong></p>\n\n<p>3) ** <a href=\"https://turi.com/products/create/docs/generated/graphlab.recommender.item_content_recommender.ItemContentRecommender.html\">Item Content Recommender</a> \nThis algorithm computes the similarity between items using the content of each item.\nThis uses  nearest neighbours and in many cases took too long. Once simple version  took about 30 min and had a CV Score of <strong>0.4617</strong> at which point I gave up.</p>\n\n<p>4) ** <a href=\"https://turi.com/products/create/docs/generated/graphlab.recommender.factorization_recommender.FactorizationRecommender.html\">Factorization Recommender</a>\nThis learns latent factors for each user and item and uses them to make predictions.</p>\n\n<p>This took about 10 min and the best CV was 0.5965</p>\n\n<p>5) Web page to Ad page similarity\nFor this I calculated the cosine similarity between the document metadata (i.e., topic, category, entity) for the web page the ads were shown in vs the web page the ad linked to, then used a boosted tree make predictions.  </p>\n\n<p>Since this required a lot of joins I used a 10% subset of the data and the best CV was 0.480</p>",
      "rawMarkdown": "I've tried a number of different approach, none of which is better than the simple Popularity Recommender first implemented by Clustifier.\r\n\r\nSince this is a large data set I'm using [GraphLab][1] and AWS spot instance with up to 32 cores and 64GB memory. This has allowed me to try the following:\r\n\r\n1) ** [Popularity Recommender][2]**\r\nThis is the same algorithm as used in BTB by Clustifier but optimised in the GL library. It's one line of code and takes less than 2 min to create a model. My CV score was **0.6442** and the submission was  0.6352 (i.e., identical to the BTB). \r\n\r\n2) ** [Popularity Recommender][3]**\r\nThis ranks items according to its similarity to other items observed for the user in question.  For this I only used the training data where clicked ==1. This took less than a min to run and the CV score was **0.5219**\r\n\r\n3) ** [Item Content Recommender][4] \r\nThis algorithm computes the similarity between items using the content of each item.\r\nThis uses  nearest neighbours and in many cases took too long. Once simple version  took about 30 min and had a CV Score of **0.4617** at which point I gave up.\r\n\r\n4) ** [Factorization Recommender][5]\r\nThis learns latent factors for each user and item and uses them to make predictions.\r\n\r\nThis took about 10 min and the best CV was 0.5965\r\n\r\n5) Web page to Ad page similarity\r\nFor this I calculated the cosine similarity between the document metadata (i.e., topic, category, entity) for the web page the ads were shown in vs the web page the ad linked to, then used a boosted tree make predictions.  \r\n\r\nSince this required a lot of joins I used a 10% subset of the data and the best CV was 0.480\r\n\r\n\r\n\r\n  [1]: https://turi.com/\r\n  [2]: https://turi.com/products/create/docs/generated/graphlab.recommender.item_similarity_recommender.ItemSimilarityRecommender.html\r\n  [3]: https://turi.com/products/create/docs/generated/graphlab.recommender.item_similarity_recommender.ItemSimilarityRecommender.html\r\n  [4]: https://turi.com/products/create/docs/generated/graphlab.recommender.item_content_recommender.ItemContentRecommender.html\r\n  [5]: https://turi.com/products/create/docs/generated/graphlab.recommender.factorization_recommender.FactorizationRecommender.html",
      "votes": null
    },
    {
      "id": "141686",
      "postDate": "10/27/2016 22:50:06",
      "content": "<p>What features are you using?</p>\n\n<p>If you are using topics, categories and entities what are you doing with the confidence numbers?</p>",
      "rawMarkdown": "What features are you using?\r\n\r\nIf you are using topics, categories and entities what are you doing with the confidence numbers?",
      "votes": null
    },
    {
      "id": "144295",
      "postDate": "11/13/2016 16:33:59",
      "content": "<p>Thank you for the report Fractal Feelings.</p>\n\n<p>May be it is worth trying a binary classification model and then order them based on the probability values to get the recommendations. It certainly helped us. </p>",
      "rawMarkdown": "Thank you for the report Fractal Feelings.\r\n\r\nMay be it is worth trying a binary classification model and then order them based on the probability values to get the recommendations. It certainly helped us.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 141686,
      "author_name": "nlothian",
      "author_url": "",
      "post_date": "10/27/2016 22:50:06",
      "content": "<p>What features are you using?</p>\n\n<p>If you are using topics, categories and entities what are you doing with the confidence numbers?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 144295,
      "author_name": "sudalairajkumar",
      "author_url": "",
      "post_date": "11/13/2016 16:33:59",
      "content": "<p>Thank you for the report Fractal Feelings.</p>\n\n<p>May be it is worth trying a binary classification model and then order them based on the probability values to get the recommendations. It certainly helped us. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "141681": "I've tried a number of different approach, none of which is better than the simple Popularity Recommender first implemented by Clustifier.\r\n\r\nSince this is a large data set I'm using [GraphLab][1] and AWS spot instance with up to 32 cores and 64GB memory. This has allowed me to try the following:\r\n\r\n1) ** [Popularity Recommender][2]**\r\nThis is the same algorithm as used in BTB by Clustifier but optimised in the GL library. It's one line of code and takes less than 2 min to create a model. My CV score was **0.6442** and the submission was  0.6352 (i.e., identical to the BTB). \r\n\r\n2) ** [Popularity Recommender][3]**\r\nThis ranks items according to its similarity to other items observed for the user in question.  For this I only used the training data where clicked ==1. This took less than a min to run and the CV score was **0.5219**\r\n\r\n3) ** [Item Content Recommender][4] \r\nThis algorithm computes the similarity between items using the content of each item.\r\nThis uses  nearest neighbours and in many cases took too long. Once simple version  took about 30 min and had a CV Score of **0.4617** at which point I gave up.\r\n\r\n4) ** [Factorization Recommender][5]\r\nThis learns latent factors for each user and item and uses them to make predictions.\r\n\r\nThis took about 10 min and the best CV was 0.5965\r\n\r\n5) Web page to Ad page similarity\r\nFor this I calculated the cosine similarity between the document metadata (i.e., topic, category, entity) for the web page the ads were shown in vs the web page the ad linked to, then used a boosted tree make predictions.  \r\n\r\nSince this required a lot of joins I used a 10% subset of the data and the best CV was 0.480\r\n\r\n\r\n\r\n  [1]: https://turi.com/\r\n  [2]: https://turi.com/products/create/docs/generated/graphlab.recommender.item_similarity_recommender.ItemSimilarityRecommender.html\r\n  [3]: https://turi.com/products/create/docs/generated/graphlab.recommender.item_similarity_recommender.ItemSimilarityRecommender.html\r\n  [4]: https://turi.com/products/create/docs/generated/graphlab.recommender.item_content_recommender.ItemContentRecommender.html\r\n  [5]: https://turi.com/products/create/docs/generated/graphlab.recommender.factorization_recommender.FactorizationRecommender.html",
    "141686": "What features are you using?\r\n\r\nIf you are using topics, categories and entities what are you doing with the confidence numbers?",
    "144295": "Thank you for the report Fractal Feelings.\r\n\r\nMay be it is worth trying a binary classification model and then order them based on the probability values to get the recommendations. It certainly helped us."
  },
  "source": "meta"
}