{
  "id": 367282,
  "title": "A Case Study - How Youtube approaches the Candidate Generation and Ranking problems?",
  "url": "/competitions/otto-recommender-system/discussion/367282",
  "author_name": "",
  "post_date": "2022-11-20T05:39:36.030567800Z",
  "votes": 22,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>As shared multiple times before in the discussions and notebooks (<a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721\" target=\"_blank\">1</a>, <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575\" target=\"_blank\">2</a>), candidate generation and ranking are two important steps in the recommendation pipeline where the number of items can be level of 100s of millions.</p>\n<p>So I wanted to share how Youtube approaches those challenges:</p>\n<p><strong>Source:</strong> <a href=\"https://dl.acm.org/doi/pdf/10.1145/2959100.2959190\" target=\"_blank\">Deep Neural Networks for YouTube Recommendations</a></p>\n<h3>System Overview</h3>\n<p>The overall structure of the system can be found in the picture below (source: Figure 2 in the paper). The system has two different neural networks: one for candidate generation and one for ranking.</p>\n<p><img src=\"https://raw.githubusercontent.com/snnclsr/kaggle_images/main/youtube_system_overview.png\" alt=\"\"></p>\n<p>The candidate generation network takes events from the user’s YouTube activity history as input and retrieves a small subset (hundreds) of videos from a large corpus. These candidates are intended to be generally relevant to the user<br>\nwith high precision. The candidate generation network only provides broad personalization via collaborative filtering.<br>\nThe similarity between users is expressed in terms of coarse features such as IDs of video watches, search query tokens and demographics.</p>\n<p>Presenting a few “best” recommendations in a list requires a fine-level representation to distinguish relative importance<br>\namong candidates with high recall. The ranking network accomplishes this task by assigning a score to each video according to a desired objective function using a rich set of features describing the video and user. The highest scoring<br>\nvideos are presented to the user, ranked by their score.</p>\n<h4>Candidate Generation</h4>\n<p>During candidate generation, the enormous YouTube corpus is winnowed down to hundreds of videos that may be<br>\nrelevant to the user. The predecessor to the recommender described here was a matrix factorization approach trained<br>\nunder rank loss.</p>\n<h5>Recommendation as Classification</h5>\n<p>The authors pose recommendation as extreme multiclass classification where the prediction problem becomes accurately classifying a specific video watch w_t at time t among millions of videos i (classes) from a corpus V based on a user U and<br>\ncontext C,<br>\n$$<br>\nP(w_t = i|U, C) = \\frac{e^{v_iu}}{\\sum_{j \\in V}{e^v_ju}}<br>\n$$<br>\nwhere u ∈ R^N represents a high-dimensional “embedding” of the user, context pair and the v_j ∈ R^N represent embeddings of each candidate video.</p>\n<p>The authors learn high dimensional embeddings for each video in a fixed vocabulary and feed these embeddings into a feedforward neural network. A user’s watch history is represented by a variable-length sequence of sparse video IDs which is mapped to a dense vector representation via the embeddings. The network requires fixed-sized dense inputs and simply averaging the embeddings performed best among several strategies (sum, component-wise max, etc.). Importantly, the embeddings are learned jointly with all other model parameters through normal gradient descent backpropagation updates. Features are concatenated into a wide first layer, followed by several layers of fully connected Rectified Linear Units (ReLU).</p>\n<p><img src=\"https://raw.githubusercontent.com/snnclsr/kaggle_images/main/youtube_candidate_generation.png\" alt=\"\"></p>\n<p>A key advantage of using deep neural networks as a generalization of matrix factorization is that arbitrary continuous<br>\nand categorical features can be easily added to the model. Search history is treated similarly to watch history - each query is tokenized into unigrams and bigrams and each token is embedded. Once averaged, the user’s tokenized, embedded queries represent a summarized dense search history. Demographic features are important for providing priors so that the recommendations behave reasonably for new users. The user’s geographic region and device are embedded and concatenated. Simple binary and continuous features such as the user’s gender, logged-in state and age are input directly into the network as real values normalized to [0, 1].</p>\n<h4>RANKING</h4>\n<p>The primary role of ranking is to use impression data to specialize and calibrate candidate predictions for the particular user interface. The authors use a deep neural network with similar architecture as candidate generation to assign an independent score to each video impression using logistic regression (below image). The list of videos is then sorted by this score and returned to the user.</p>\n<p><img src=\"https://raw.githubusercontent.com/snnclsr/kaggle_images/main/youtube_ranking.png\" alt=\"\"></p>",
  "messages": [
    {
      "id": "2036715",
      "postDate": "11/20/2022 05:39:36",
      "content": "<p>Hi everyone,</p>\n<p>As shared multiple times before in the discussions and notebooks (<a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721\" target=\"_blank\">1</a>, <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575\" target=\"_blank\">2</a>), candidate generation and ranking are two important steps in the recommendation pipeline where the number of items can be level of 100s of millions.</p>\n<p>So I wanted to share how Youtube approaches those challenges:</p>\n<p><strong>Source:</strong> <a href=\"https://dl.acm.org/doi/pdf/10.1145/2959100.2959190\" target=\"_blank\">Deep Neural Networks for YouTube Recommendations</a></p>\n<h3>System Overview</h3>\n<p>The overall structure of the system can be found in the picture below (source: Figure 2 in the paper). The system has two different neural networks: one for candidate generation and one for ranking.</p>\n<p><img src=\"https://raw.githubusercontent.com/snnclsr/kaggle_images/main/youtube_system_overview.png\" alt=\"\"></p>\n<p>The candidate generation network takes events from the user’s YouTube activity history as input and retrieves a small subset (hundreds) of videos from a large corpus. These candidates are intended to be generally relevant to the user<br>\nwith high precision. The candidate generation network only provides broad personalization via collaborative filtering.<br>\nThe similarity between users is expressed in terms of coarse features such as IDs of video watches, search query tokens and demographics.</p>\n<p>Presenting a few “best” recommendations in a list requires a fine-level representation to distinguish relative importance<br>\namong candidates with high recall. The ranking network accomplishes this task by assigning a score to each video according to a desired objective function using a rich set of features describing the video and user. The highest scoring<br>\nvideos are presented to the user, ranked by their score.</p>\n<h4>Candidate Generation</h4>\n<p>During candidate generation, the enormous YouTube corpus is winnowed down to hundreds of videos that may be<br>\nrelevant to the user. The predecessor to the recommender described here was a matrix factorization approach trained<br>\nunder rank loss.</p>\n<h5>Recommendation as Classification</h5>\n<p>The authors pose recommendation as extreme multiclass classification where the prediction problem becomes accurately classifying a specific video watch w_t at time t among millions of videos i (classes) from a corpus V based on a user U and<br>\ncontext C,<br>\n$$<br>\nP(w_t = i|U, C) = \\frac{e^{v_iu}}{\\sum_{j \\in V}{e^v_ju}}<br>\n$$<br>\nwhere u ∈ R^N represents a high-dimensional “embedding” of the user, context pair and the v_j ∈ R^N represent embeddings of each candidate video.</p>\n<p>The authors learn high dimensional embeddings for each video in a fixed vocabulary and feed these embeddings into a feedforward neural network. A user’s watch history is represented by a variable-length sequence of sparse video IDs which is mapped to a dense vector representation via the embeddings. The network requires fixed-sized dense inputs and simply averaging the embeddings performed best among several strategies (sum, component-wise max, etc.). Importantly, the embeddings are learned jointly with all other model parameters through normal gradient descent backpropagation updates. Features are concatenated into a wide first layer, followed by several layers of fully connected Rectified Linear Units (ReLU).</p>\n<p><img src=\"https://raw.githubusercontent.com/snnclsr/kaggle_images/main/youtube_candidate_generation.png\" alt=\"\"></p>\n<p>A key advantage of using deep neural networks as a generalization of matrix factorization is that arbitrary continuous<br>\nand categorical features can be easily added to the model. Search history is treated similarly to watch history - each query is tokenized into unigrams and bigrams and each token is embedded. Once averaged, the user’s tokenized, embedded queries represent a summarized dense search history. Demographic features are important for providing priors so that the recommendations behave reasonably for new users. The user’s geographic region and device are embedded and concatenated. Simple binary and continuous features such as the user’s gender, logged-in state and age are input directly into the network as real values normalized to [0, 1].</p>\n<h4>RANKING</h4>\n<p>The primary role of ranking is to use impression data to specialize and calibrate candidate predictions for the particular user interface. The authors use a deep neural network with similar architecture as candidate generation to assign an independent score to each video impression using logistic regression (below image). The list of videos is then sorted by this score and returned to the user.</p>\n<p><img src=\"https://raw.githubusercontent.com/snnclsr/kaggle_images/main/youtube_ranking.png\" alt=\"\"></p>",
      "rawMarkdown": "Hi everyone,\n\nAs shared multiple times before in the discussions and notebooks ([1](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721), [2](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575)), candidate generation and ranking are two important steps in the recommendation pipeline where the number of items can be level of 100s of millions.\n\nSo I wanted to share how Youtube approaches those challenges:\n\n**Source:** [Deep Neural Networks for YouTube Recommendations](https://dl.acm.org/doi/pdf/10.1145/2959100.2959190)\n\n### System Overview\nThe overall structure of the system can be found in the picture below (source: Figure 2 in the paper). The system has two different neural networks: one for candidate generation and one for ranking.\n\n![](https://raw.githubusercontent.com/snnclsr/kaggle_images/main/youtube_system_overview.png)\n\nThe candidate generation network takes events from the user’s YouTube activity history as input and retrieves a small subset (hundreds) of videos from a large corpus. These candidates are intended to be generally relevant to the user\nwith high precision. The candidate generation network only provides broad personalization via collaborative filtering.\nThe similarity between users is expressed in terms of coarse features such as IDs of video watches, search query tokens and demographics.\n\nPresenting a few “best” recommendations in a list requires a fine-level representation to distinguish relative importance\namong candidates with high recall. The ranking network accomplishes this task by assigning a score to each video according to a desired objective function using a rich set of features describing the video and user. The highest scoring\nvideos are presented to the user, ranked by their score.\n\n#### Candidate Generation\nDuring candidate generation, the enormous YouTube corpus is winnowed down to hundreds of videos that may be\nrelevant to the user. The predecessor to the recommender described here was a matrix factorization approach trained\nunder rank loss.\n\n##### Recommendation as Classification\n\nThe authors pose recommendation as extreme multiclass classification where the prediction problem becomes accurately classifying a specific video watch w_t at time t among millions of videos i (classes) from a corpus V based on a user U and\ncontext C,\n$$\nP(w_t = i|U, C) = \\frac{e^{v_iu}}{\\sum_{j \\in V}{e^v_ju}}\n$$\nwhere u ∈ R^N represents a high-dimensional “embedding” of the user, context pair and the v_j ∈ R^N represent embeddings of each candidate video.\n\nThe authors learn high dimensional embeddings for each video in a fixed vocabulary and feed these embeddings into a feedforward neural network. A user’s watch history is represented by a variable-length sequence of sparse video IDs which is mapped to a dense vector representation via the embeddings. The network requires fixed-sized dense inputs and simply averaging the embeddings performed best among several strategies (sum, component-wise max, etc.). Importantly, the embeddings are learned jointly with all other model parameters through normal gradient descent backpropagation updates. Features are concatenated into a wide first layer, followed by several layers of fully connected Rectified Linear Units (ReLU).\n\n![](https://raw.githubusercontent.com/snnclsr/kaggle_images/main/youtube_candidate_generation.png)\n\nA key advantage of using deep neural networks as a generalization of matrix factorization is that arbitrary continuous\nand categorical features can be easily added to the model. Search history is treated similarly to watch history - each query is tokenized into unigrams and bigrams and each token is embedded. Once averaged, the user’s tokenized, embedded queries represent a summarized dense search history. Demographic features are important for providing priors so that the recommendations behave reasonably for new users. The user’s geographic region and device are embedded and concatenated. Simple binary and continuous features such as the user’s gender, logged-in state and age are input directly into the network as real values normalized to [0, 1].\n\n#### RANKING\n\nThe primary role of ranking is to use impression data to specialize and calibrate candidate predictions for the particular user interface. The authors use a deep neural network with similar architecture as candidate generation to assign an independent score to each video impression using logistic regression (below image). The list of videos is then sorted by this score and returned to the user.\n\n![](https://raw.githubusercontent.com/snnclsr/kaggle_images/main/youtube_ranking.png)",
      "votes": null
    },
    {
      "id": "2037685",
      "postDate": "11/20/2022 20:19:26",
      "content": "<p>There are differences, in your use case the order of research can be have a influence. This differente with youtube ranking.</p>",
      "rawMarkdown": "There are differences, in your use case the order of research can be have a influence. This differente with youtube ranking.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2037685,
      "author_name": "lgregory",
      "author_url": "",
      "post_date": "11/20/2022 20:19:26",
      "content": "<p>There are differences, in your use case the order of research can be have a influence. This differente with youtube ranking.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2036715": "Hi everyone,\n\nAs shared multiple times before in the discussions and notebooks ([1](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721), [2](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575)), candidate generation and ranking are two important steps in the recommendation pipeline where the number of items can be level of 100s of millions.\n\nSo I wanted to share how Youtube approaches those challenges:\n\n**Source:** [Deep Neural Networks for YouTube Recommendations](https://dl.acm.org/doi/pdf/10.1145/2959100.2959190)\n\n### System Overview\nThe overall structure of the system can be found in the picture below (source: Figure 2 in the paper). The system has two different neural networks: one for candidate generation and one for ranking.\n\n![](https://raw.githubusercontent.com/snnclsr/kaggle_images/main/youtube_system_overview.png)\n\nThe candidate generation network takes events from the user’s YouTube activity history as input and retrieves a small subset (hundreds) of videos from a large corpus. These candidates are intended to be generally relevant to the user\nwith high precision. The candidate generation network only provides broad personalization via collaborative filtering.\nThe similarity between users is expressed in terms of coarse features such as IDs of video watches, search query tokens and demographics.\n\nPresenting a few “best” recommendations in a list requires a fine-level representation to distinguish relative importance\namong candidates with high recall. The ranking network accomplishes this task by assigning a score to each video according to a desired objective function using a rich set of features describing the video and user. The highest scoring\nvideos are presented to the user, ranked by their score.\n\n#### Candidate Generation\nDuring candidate generation, the enormous YouTube corpus is winnowed down to hundreds of videos that may be\nrelevant to the user. The predecessor to the recommender described here was a matrix factorization approach trained\nunder rank loss.\n\n##### Recommendation as Classification\n\nThe authors pose recommendation as extreme multiclass classification where the prediction problem becomes accurately classifying a specific video watch w_t at time t among millions of videos i (classes) from a corpus V based on a user U and\ncontext C,\n$$\nP(w_t = i|U, C) = \\frac{e^{v_iu}}{\\sum_{j \\in V}{e^v_ju}}\n$$\nwhere u ∈ R^N represents a high-dimensional “embedding” of the user, context pair and the v_j ∈ R^N represent embeddings of each candidate video.\n\nThe authors learn high dimensional embeddings for each video in a fixed vocabulary and feed these embeddings into a feedforward neural network. A user’s watch history is represented by a variable-length sequence of sparse video IDs which is mapped to a dense vector representation via the embeddings. The network requires fixed-sized dense inputs and simply averaging the embeddings performed best among several strategies (sum, component-wise max, etc.). Importantly, the embeddings are learned jointly with all other model parameters through normal gradient descent backpropagation updates. Features are concatenated into a wide first layer, followed by several layers of fully connected Rectified Linear Units (ReLU).\n\n![](https://raw.githubusercontent.com/snnclsr/kaggle_images/main/youtube_candidate_generation.png)\n\nA key advantage of using deep neural networks as a generalization of matrix factorization is that arbitrary continuous\nand categorical features can be easily added to the model. Search history is treated similarly to watch history - each query is tokenized into unigrams and bigrams and each token is embedded. Once averaged, the user’s tokenized, embedded queries represent a summarized dense search history. Demographic features are important for providing priors so that the recommendations behave reasonably for new users. The user’s geographic region and device are embedded and concatenated. Simple binary and continuous features such as the user’s gender, logged-in state and age are input directly into the network as real values normalized to [0, 1].\n\n#### RANKING\n\nThe primary role of ranking is to use impression data to specialize and calibrate candidate predictions for the particular user interface. The authors use a deep neural network with similar architecture as candidate generation to assign an independent score to each video impression using logistic regression (below image). The list of videos is then sorted by this score and returned to the user.\n\n![](https://raw.githubusercontent.com/snnclsr/kaggle_images/main/youtube_ranking.png)",
    "2037685": "There are differences, in your use case the order of research can be have a influence. This differente with youtube ranking."
  },
  "source": "meta"
}