{
  "id": 383260,
  "title": "Proof of concept vanilla \"Markov Decision Process\"",
  "url": "/competitions/otto-recommender-system/discussion/383260",
  "author_name": "Hakan Yilmazer",
  "post_date": "2023-02-02T22:11:52.380000",
  "votes": 4,
  "comment_count": 0,
  "views": 0,
  "content": "<p>First of all, I would like to thank the Otto team <a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a> and Kaggle team for this competition.</p>\n<p>The one of the problem for researchers in Session-based recommendation systems is to find suitable evaluation datasets. Mostly benchmark datasets are common dataset which is used more general problems. That's why I think this dataset is so valuable for future researches.<br>\nI'm sure that it would be label as 'well-known' datasets for researchers in literature.</p>\n<p>It was a first but very enjoyable and exciting kaggling adventure for me.<br>\nI have made a lot of profits.<br>\nObserving how the well-known Recommendation Models developed are evaluated with big data inspired my future algorithm/model ideas.<br>\nSeeing how the dataset and RS models are implemented by Kaggle Competitors has been good feedback on the pros and cons of these models.<br>\nAlso, Polars is now inevitable.</p>\n<p>A beautiful community was formed here during the entire competition. Many thanks to everyone who participated, offered ideas and added value.<br>\nI would like to special thank <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> and <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>, who helped me adapt to the competition and create a bootstrap with their ideas, advice and codes.</p>\n<p><strong>Candidate Generation with Markov Decision Process</strong></p>\n<p>In this competition, I focused on the idea of how to increase the 'Recall' score effectively with a efficiency 'single-model'. Because no matter which boosting model you use, you are as strong as the intersection value of your candidate items with ground-truth items. Therefore, we see that the models that are successful in the leaderboard use more than one candidate item generation algorithm.</p>\n<p>One of the points I focused on was to produce a 'Multi-Objective' candidate item with a single model. In other words, instead of determining the next item (Item CF, WordVec, etc.), I focused on which action will make the next click.</p>\n<p>This goal led me to the idea of designing a simplified 'Markov Decision Process' model.<br>\nMarkov Decision Process, a stochastic decision-making process that uses a mathematical framework to model the decision-making of a dynamic system. It is used in scenarios where the results are either random or controlled by a decision maker, which makes sequential decisions over time.</p>\n<p>MDP is one of the leading models used in Session-based systems. There are successful studies on this subject [see ref 1, 2].</p>\n<p>In Markov Decision Processes, you try to predict not only the next item but also the action. The choices you make in the next step in Markov Decision Processes also affect the future choices. I matched it to the life. The decisions we make affect our future decisions. I also think that this model is very suitable for this competition data. Because the data we have is a real data. This data is user interactions and clicks. The user clicks on a product on the website, adds the product to the cart on the product page, can place an order, or click on similar, trending products on the side of the page or below. In addition, while on the product page, you can search in the search box and switch to another product page. Each click spreads the probabilities for the user.</p>\n<p><img src=\"https://raw.githubusercontent.com/otto-de/recsys-dataset/main/.readme/ground_truth.png\" alt=\"Ground Truth Schema\"></p>\n<p>While simulating the behavior of a sample user in the diagram above, different or same products can be selected from different actions and we were trying to select this user's target items and the event types of the items. This scheme is very similar to the undirected graph scheme associated with MDP.</p>\n<p><strong>Basic Concepts of Markov Decision Process</strong></p>\n<p>A MDP consists of the following elements:</p>\n<ol>\n<li><em>T</em> is all decision times.</li>\n<li><em>S</em> is a set of states, which is a set of all possible states of the system.</li>\n<li><em>A</em> is set of actions</li>\n<li><em>P</em> is set of Transition Probabilities</li>\n<li><em>R</em> is a Reward Function</li>\n</ol>\n<p><img src=\"https://upload.wikimedia.org/wikipedia/commons/thumb/a/ad/Markov_Decision_Process.svg/400px-Markov_Decision_Process.svg.png\" alt=\"Image from Wikipedi\"> <br>\nImage Source: <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3345899%2Ff0ab35dd8b065fe4b151be2d4a30f61e%2FMarkov_Decision_Process.svg.png?generation=1675375332680029&amp;alt=media\" alt=\"\"><a href=\"https://en.wikipedia.org/wiki/Markov_decision_process\" target=\"_blank\">https://en.wikipedia.org/wiki/Markov_decision_process</a></p>\n<p>In this competition manner; </p>\n<ul>\n<li>S is candidate item set,</li>\n<li>A is set of actions those are {'clicks', 'carts', 'orders'}, </li>\n<li>P is the probability matrix which one of its value is the switching probability from an item to next item with a defined action. </li>\n<li>R is reward of a probable transition.</li>\n</ul>\n<p><strong>Basic Approach</strong></p>\n<ul>\n<li>I used <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> optimized cross-validation dataset for train and evaluation</li>\n<li>I trained with Train and Test Set.</li>\n<li>Validation with Valid Set</li>\n<li>CV Score: 562</li>\n<li>LB Score: 572</li>\n</ul>\n<p>Detailed explanation of codes are in the below Notebook:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/fotoizzet/proof-of-concept-markov-decision-process\" target=\"_blank\">https://www.kaggle.com/code/fotoizzet/proof-of-concept-markov-decision-process</a></li>\n</ul>\n<h3>References</h3>\n<ol>\n<li>An MDP-based Recommender System, <a href=\"https://arxiv.org/abs/1301.0600\" target=\"_blank\">https://arxiv.org/abs/1301.0600</a></li>\n<li>Factorizing Personalized Markov Chains for Next-Basket Recommendation, <a href=\"https://www.ra.ethz.ch/cdstore/www2010/www/p811.pdf\" target=\"_blank\">https://www.ra.ethz.ch/cdstore/www2010/www/p811.pdf</a></li>\n</ol>",
  "messages": [
    {
      "id": 2127438,
      "postDate": "2023-02-02T22:11:52.380Z",
      "content": "<p>First of all, I would like to thank the Otto team <a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a> and Kaggle team for this competition.</p>\n<p>The one of the problem for researchers in Session-based recommendation systems is to find suitable evaluation datasets. Mostly benchmark datasets are common dataset which is used more general problems. That's why I think this dataset is so valuable for future researches.<br>\nI'm sure that it would be label as 'well-known' datasets for researchers in literature.</p>\n<p>It was a first but very enjoyable and exciting kaggling adventure for me.<br>\nI have made a lot of profits.<br>\nObserving how the well-known Recommendation Models developed are evaluated with big data inspired my future algorithm/model ideas.<br>\nSeeing how the dataset and RS models are implemented by Kaggle Competitors has been good feedback on the pros and cons of these models.<br>\nAlso, Polars is now inevitable.</p>\n<p>A beautiful community was formed here during the entire competition. Many thanks to everyone who participated, offered ideas and added value.<br>\nI would like to special thank <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> and <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>, who helped me adapt to the competition and create a bootstrap with their ideas, advice and codes.</p>\n<p><strong>Candidate Generation with Markov Decision Process</strong></p>\n<p>In this competition, I focused on the idea of how to increase the 'Recall' score effectively with a efficiency 'single-model'. Because no matter which boosting model you use, you are as strong as the intersection value of your candidate items with ground-truth items. Therefore, we see that the models that are successful in the leaderboard use more than one candidate item generation algorithm.</p>\n<p>One of the points I focused on was to produce a 'Multi-Objective' candidate item with a single model. In other words, instead of determining the next item (Item CF, WordVec, etc.), I focused on which action will make the next click.</p>\n<p>This goal led me to the idea of designing a simplified 'Markov Decision Process' model.<br>\nMarkov Decision Process, a stochastic decision-making process that uses a mathematical framework to model the decision-making of a dynamic system. It is used in scenarios where the results are either random or controlled by a decision maker, which makes sequential decisions over time.</p>\n<p>MDP is one of the leading models used in Session-based systems. There are successful studies on this subject [see ref 1, 2].</p>\n<p>In Markov Decision Processes, you try to predict not only the next item but also the action. The choices you make in the next step in Markov Decision Processes also affect the future choices. I matched it to the life. The decisions we make affect our future decisions. I also think that this model is very suitable for this competition data. Because the data we have is a real data. This data is user interactions and clicks. The user clicks on a product on the website, adds the product to the cart on the product page, can place an order, or click on similar, trending products on the side of the page or below. In addition, while on the product page, you can search in the search box and switch to another product page. Each click spreads the probabilities for the user.</p>\n<p><img src=\"https://raw.githubusercontent.com/otto-de/recsys-dataset/main/.readme/ground_truth.png\" alt=\"Ground Truth Schema\"></p>\n<p>While simulating the behavior of a sample user in the diagram above, different or same products can be selected from different actions and we were trying to select this user's target items and the event types of the items. This scheme is very similar to the undirected graph scheme associated with MDP.</p>\n<p><strong>Basic Concepts of Markov Decision Process</strong></p>\n<p>A MDP consists of the following elements:</p>\n<ol>\n<li><em>T</em> is all decision times.</li>\n<li><em>S</em> is a set of states, which is a set of all possible states of the system.</li>\n<li><em>A</em> is set of actions</li>\n<li><em>P</em> is set of Transition Probabilities</li>\n<li><em>R</em> is a Reward Function</li>\n</ol>\n<p><img src=\"https://upload.wikimedia.org/wikipedia/commons/thumb/a/ad/Markov_Decision_Process.svg/400px-Markov_Decision_Process.svg.png\" alt=\"Image from Wikipedi\"> <br>\nImage Source: <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3345899%2Ff0ab35dd8b065fe4b151be2d4a30f61e%2FMarkov_Decision_Process.svg.png?generation=1675375332680029&amp;alt=media\" alt=\"\"><a href=\"https://en.wikipedia.org/wiki/Markov_decision_process\" target=\"_blank\">https://en.wikipedia.org/wiki/Markov_decision_process</a></p>\n<p>In this competition manner; </p>\n<ul>\n<li>S is candidate item set,</li>\n<li>A is set of actions those are {'clicks', 'carts', 'orders'}, </li>\n<li>P is the probability matrix which one of its value is the switching probability from an item to next item with a defined action. </li>\n<li>R is reward of a probable transition.</li>\n</ul>\n<p><strong>Basic Approach</strong></p>\n<ul>\n<li>I used <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> optimized cross-validation dataset for train and evaluation</li>\n<li>I trained with Train and Test Set.</li>\n<li>Validation with Valid Set</li>\n<li>CV Score: 562</li>\n<li>LB Score: 572</li>\n</ul>\n<p>Detailed explanation of codes are in the below Notebook:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/fotoizzet/proof-of-concept-markov-decision-process\" target=\"_blank\">https://www.kaggle.com/code/fotoizzet/proof-of-concept-markov-decision-process</a></li>\n</ul>\n<h3>References</h3>\n<ol>\n<li>An MDP-based Recommender System, <a href=\"https://arxiv.org/abs/1301.0600\" target=\"_blank\">https://arxiv.org/abs/1301.0600</a></li>\n<li>Factorizing Personalized Markov Chains for Next-Basket Recommendation, <a href=\"https://www.ra.ethz.ch/cdstore/www2010/www/p811.pdf\" target=\"_blank\">https://www.ra.ethz.ch/cdstore/www2010/www/p811.pdf</a></li>\n</ol>",
      "rawMarkdown": "First of all, I would like to thank the Otto team @pnormann and Kaggle team for this competition.\n\nThe one of the problem for researchers in Session-based recommendation systems is to find suitable evaluation datasets. Mostly benchmark datasets are common dataset which is used more general problems. That's why I think this dataset is so valuable for future researches.\nI'm sure that it would be label as 'well-known' datasets for researchers in literature.\n\nIt was a first but very enjoyable and exciting kaggling adventure for me.\nI have made a lot of profits.\nObserving how the well-known Recommendation Models developed are evaluated with big data inspired my future algorithm/model ideas.\nSeeing how the dataset and RS models are implemented by Kaggle Competitors has been good feedback on the pros and cons of these models.\nAlso, Polars is now inevitable.\n\nA beautiful community was formed here during the entire competition. Many thanks to everyone who participated, offered ideas and added value.\nI would like to special thank @cdeotte and @radek1, who helped me adapt to the competition and create a bootstrap with their ideas, advice and codes.\n\n**Candidate Generation with Markov Decision Process**\n\nIn this competition, I focused on the idea of how to increase the 'Recall' score effectively with a efficiency 'single-model'. Because no matter which boosting model you use, you are as strong as the intersection value of your candidate items with ground-truth items. Therefore, we see that the models that are successful in the leaderboard use more than one candidate item generation algorithm.\n\nOne of the points I focused on was to produce a 'Multi-Objective' candidate item with a single model. In other words, instead of determining the next item (Item CF, WordVec, etc.), I focused on which action will make the next click.\n\nThis goal led me to the idea of designing a simplified 'Markov Decision Process' model.\nMarkov Decision Process, a stochastic decision-making process that uses a mathematical framework to model the decision-making of a dynamic system. It is used in scenarios where the results are either random or controlled by a decision maker, which makes sequential decisions over time.\n\nMDP is one of the leading models used in Session-based systems. There are successful studies on this subject [see ref 1, 2].\n\nIn Markov Decision Processes, you try to predict not only the next item but also the action. The choices you make in the next step in Markov Decision Processes also affect the future choices. I matched it to the life. The decisions we make affect our future decisions. I also think that this model is very suitable for this competition data. Because the data we have is a real data. This data is user interactions and clicks. The user clicks on a product on the website, adds the product to the cart on the product page, can place an order, or click on similar, trending products on the side of the page or below. In addition, while on the product page, you can search in the search box and switch to another product page. Each click spreads the probabilities for the user.\n\n![Ground Truth Schema](https://raw.githubusercontent.com/otto-de/recsys-dataset/main/.readme/ground_truth.png)\n\nWhile simulating the behavior of a sample user in the diagram above, different or same products can be selected from different actions and we were trying to select this user's target items and the event types of the items. This scheme is very similar to the undirected graph scheme associated with MDP.\n\n**Basic Concepts of Markov Decision Process**\n\nA MDP consists of the following elements:\n1. *T* is all decision times.\n2. *S* is a set of states, which is a set of all possible states of the system.\n3. *A* is set of actions\n4. *P* is set of Transition Probabilities\n5. *R* is a Reward Function\n\n![Image from Wikipedi](https://upload.wikimedia.org/wikipedia/commons/thumb/a/ad/Markov_Decision_Process.svg/400px-Markov_Decision_Process.svg.png) \nImage Source: ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3345899%2Ff0ab35dd8b065fe4b151be2d4a30f61e%2FMarkov_Decision_Process.svg.png?generation=1675375332680029&alt=media)https://en.wikipedia.org/wiki/Markov_decision_process\n\nIn this competition manner; \n*     S is candidate item set,\n*     A is set of actions those are {'clicks', 'carts', 'orders'}, \n*     P is the probability matrix which one of its value is the switching probability from an item to next item with a defined action. \n*     R is reward of a probable transition.\n\n**Basic Approach**\n\n* I used @radek1 optimized cross-validation dataset for train and evaluation\n* I trained with Train and Test Set.\n* Validation with Valid Set\n* CV Score: 562\n* LB Score: 572\n\nDetailed explanation of codes are in the below Notebook:\n\n* https://www.kaggle.com/code/fotoizzet/proof-of-concept-markov-decision-process\n\n### References\n1.  An MDP-based Recommender System, https://arxiv.org/abs/1301.0600\n2.  Factorizing Personalized Markov Chains for Next-Basket Recommendation, https://www.ra.ethz.ch/cdstore/www2010/www/p811.pdf",
      "votes": 4
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2127438": "First of all, I would like to thank the Otto team @pnormann and Kaggle team for this competition.\n\nThe one of the problem for researchers in Session-based recommendation systems is to find suitable evaluation datasets. Mostly benchmark datasets are common dataset which is used more general problems. That's why I think this dataset is so valuable for future researches.\nI'm sure that it would be label as 'well-known' datasets for researchers in literature.\n\nIt was a first but very enjoyable and exciting kaggling adventure for me.\nI have made a lot of profits.\nObserving how the well-known Recommendation Models developed are evaluated with big data inspired my future algorithm/model ideas.\nSeeing how the dataset and RS models are implemented by Kaggle Competitors has been good feedback on the pros and cons of these models.\nAlso, Polars is now inevitable.\n\nA beautiful community was formed here during the entire competition. Many thanks to everyone who participated, offered ideas and added value.\nI would like to special thank @cdeotte and @radek1, who helped me adapt to the competition and create a bootstrap with their ideas, advice and codes.\n\n**Candidate Generation with Markov Decision Process**\n\nIn this competition, I focused on the idea of how to increase the 'Recall' score effectively with a efficiency 'single-model'. Because no matter which boosting model you use, you are as strong as the intersection value of your candidate items with ground-truth items. Therefore, we see that the models that are successful in the leaderboard use more than one candidate item generation algorithm.\n\nOne of the points I focused on was to produce a 'Multi-Objective' candidate item with a single model. In other words, instead of determining the next item (Item CF, WordVec, etc.), I focused on which action will make the next click.\n\nThis goal led me to the idea of designing a simplified 'Markov Decision Process' model.\nMarkov Decision Process, a stochastic decision-making process that uses a mathematical framework to model the decision-making of a dynamic system. It is used in scenarios where the results are either random or controlled by a decision maker, which makes sequential decisions over time.\n\nMDP is one of the leading models used in Session-based systems. There are successful studies on this subject [see ref 1, 2].\n\nIn Markov Decision Processes, you try to predict not only the next item but also the action. The choices you make in the next step in Markov Decision Processes also affect the future choices. I matched it to the life. The decisions we make affect our future decisions. I also think that this model is very suitable for this competition data. Because the data we have is a real data. This data is user interactions and clicks. The user clicks on a product on the website, adds the product to the cart on the product page, can place an order, or click on similar, trending products on the side of the page or below. In addition, while on the product page, you can search in the search box and switch to another product page. Each click spreads the probabilities for the user.\n\n![Ground Truth Schema](https://raw.githubusercontent.com/otto-de/recsys-dataset/main/.readme/ground_truth.png)\n\nWhile simulating the behavior of a sample user in the diagram above, different or same products can be selected from different actions and we were trying to select this user's target items and the event types of the items. This scheme is very similar to the undirected graph scheme associated with MDP.\n\n**Basic Concepts of Markov Decision Process**\n\nA MDP consists of the following elements:\n1. *T* is all decision times.\n2. *S* is a set of states, which is a set of all possible states of the system.\n3. *A* is set of actions\n4. *P* is set of Transition Probabilities\n5. *R* is a Reward Function\n\n![Image from Wikipedi](https://upload.wikimedia.org/wikipedia/commons/thumb/a/ad/Markov_Decision_Process.svg/400px-Markov_Decision_Process.svg.png) \nImage Source: ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3345899%2Ff0ab35dd8b065fe4b151be2d4a30f61e%2FMarkov_Decision_Process.svg.png?generation=1675375332680029&alt=media)https://en.wikipedia.org/wiki/Markov_decision_process\n\nIn this competition manner; \n*     S is candidate item set,\n*     A is set of actions those are {'clicks', 'carts', 'orders'}, \n*     P is the probability matrix which one of its value is the switching probability from an item to next item with a defined action. \n*     R is reward of a probable transition.\n\n**Basic Approach**\n\n* I used @radek1 optimized cross-validation dataset for train and evaluation\n* I trained with Train and Test Set.\n* Validation with Valid Set\n* CV Score: 562\n* LB Score: 572\n\nDetailed explanation of codes are in the below Notebook:\n\n* https://www.kaggle.com/code/fotoizzet/proof-of-concept-markov-decision-process\n\n### References\n1.  An MDP-based Recommender System, https://arxiv.org/abs/1301.0600\n2.  Factorizing Personalized Markov Chains for Next-Basket Recommendation, https://www.ra.ethz.ch/cdstore/www2010/www/p811.pdf"
  }
}