{
  "id": 58948,
  "title": "Deal Probability Strange Findings",
  "url": "/competitions/avito-demand-prediction/discussion/58948",
  "author_name": "Ahmet Erdem",
  "post_date": "2018-06-15T19:10:12.953000",
  "votes": 30,
  "comment_count": 9,
  "views": 0,
  "content": "<p><img src=\"https://i.imgur.com/KpmdQeU.png\" alt=\"enter image description here\"></p>\n\n<p>Deal probability values are very discrete. Therefore I don't think they are simple likelihoods. There must be some parent_category related parameter in its calculation because most of the values only exist for single parent category. Also the fact that nobody explains what exactly deal probability is makes the situation even more suspicious. And it seems a global minmax scaling is applied on them.</p>",
  "messages": [
    {
      "id": 343651,
      "postDate": "2018-06-15T19:10:12.953Z",
      "content": "<p><img src=\"https://i.imgur.com/KpmdQeU.png\" alt=\"enter image description here\"></p>\n\n<p>Deal probability values are very discrete. Therefore I don't think they are simple likelihoods. There must be some parent_category related parameter in its calculation because most of the values only exist for single parent category. Also the fact that nobody explains what exactly deal probability is makes the situation even more suspicious. And it seems a global minmax scaling is applied on them.</p>",
      "rawMarkdown": "![enter image description here][1]\n\nDeal probability values are very discrete. Therefore I don't think they are simple likelihoods. There must be some parent_category related parameter in its calculation because most of the values only exist for single parent category. Also the fact that nobody explains what exactly deal probability is makes the situation even more suspicious. And it seems a global minmax scaling is applied on them.\n\n\n  [1]: https://i.imgur.com/KpmdQeU.png",
      "votes": 30
    },
    {
      "id": 344543,
      "postDate": "2018-06-18T07:23:44.867Z",
      "content": "<p>There are some more \"strange findings\" hidden inside deal probability variable and waiting to be revealed ;) </p>\n\n<p>Tried to show part of them in a small notebook here: <a href=\"https://www.kaggle.com/alijs1/target-variable-some-interesting-insights\">https://www.kaggle.com/alijs1/target-variable-some-interesting-insights</a></p>",
      "rawMarkdown": "There are some more \"strange findings\" hidden inside deal probability variable and waiting to be revealed ;) \n\nTried to show part of them in a small notebook here: https://www.kaggle.com/alijs1/target-variable-some-interesting-insights",
      "votes": 5
    },
    {
      "id": 344560,
      "postDate": "2018-06-18T08:16:01.690Z",
      "content": "<p>@alijs @AhmetErdem There is more evidence that the deal probability was generated at least partially by a rule based system.  Take a look at the deal probability distribution for each <code>param_2</code> category (<a href=\"https://www.kaggle.com/wesamelshamy/the-good-deal-feature\">plots at the bottom of this kernel</a>), you will find the following:</p>\n\n<ul>\n<li>There is a big gap between 0 and ~0.1 values</li>\n<li>About 80% of the ads have exactly 0 deal probability value</li>\n<li>As you both mentioned, the deal probability values are discrete</li>\n</ul>\n\n<p>How about training a zero deal probability model?</p>",
      "rawMarkdown": "@alijs @AhmetErdem There is more evidence that the deal probability was generated at least partially by a rule based system.  Take a look at the deal probability distribution for each `param_2` category ([plots at the bottom of this kernel][1]), you will find the following:\n\n- There is a big gap between 0 and ~0.1 values\n- About 80% of the ads have exactly 0 deal probability value\n- As you both mentioned, the deal probability values are discrete\n\nHow about training a zero deal probability model?\n\n[1]: https://www.kaggle.com/wesamelshamy/the-good-deal-feature",
      "votes": 1
    },
    {
      "id": 343799,
      "postDate": "2018-06-16T06:45:26.760Z",
      "content": "<p>len(train['deal_probability'].unique())=18407  </p>",
      "rawMarkdown": "len(train['deal_probability'].unique())=18407  ",
      "votes": 1,
      "replies": [
        {
          "id": 344563,
          "postDate": "2018-06-18T08:19:37.963Z",
          "content": "<p>More succinctly: <code>train['deal_probability'].nunique()</code></p>",
          "rawMarkdown": "More succinctly: `train['deal_probability'].nunique()`",
          "votes": 1
        }
      ]
    },
    {
      "id": 345870,
      "postDate": "2018-06-20T16:21:56.257Z",
      "content": "<p>It seems to be a reverse engineering problem rather than a prediction one :-)</p>",
      "rawMarkdown": "It seems to be a reverse engineering problem rather than a prediction one :-)",
      "votes": 2,
      "replies": [
        {
          "id": 345872,
          "postDate": "2018-06-20T16:25:55.577Z",
          "content": "<p>Mostly, yes.  We are building a model that can predict what Avito model predicts.</p>",
          "rawMarkdown": "Mostly, yes.  We are building a model that can predict what Avito model predicts.",
          "votes": 1
        },
        {
          "id": 345875,
          "postDate": "2018-06-20T16:29:18.407Z",
          "content": "<p>So, if you have a very brilliant insight, but they had not in their model, it will render useless...</p>",
          "rawMarkdown": "So, if you have a very brilliant insight, but they had not in their model, it will render useless...",
          "votes": 1
        },
        {
          "id": 345900,
          "postDate": "2018-06-20T17:26:00.533Z",
          "content": "<p>It seems to be that the features are rated before joining the pipeline for classification.</p>",
          "rawMarkdown": "It seems to be that the features are rated before joining the pipeline for classification.",
          "votes": 1
        }
      ]
    },
    {
      "id": 344398,
      "postDate": "2018-06-17T21:45:58.447Z",
      "content": "<p>Given they are pretty discrete, would it make sense to train some kind of classification models and use output as features?</p>",
      "rawMarkdown": "Given they are pretty discrete, would it make sense to train some kind of classification models and use output as features?",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 344543,
      "author_name": "alijs",
      "author_url": "",
      "post_date": "2018-06-18T07:23:44.867000",
      "content": "<p>There are some more \"strange findings\" hidden inside deal probability variable and waiting to be revealed ;) </p>\n\n<p>Tried to show part of them in a small notebook here: <a href=\"https://www.kaggle.com/alijs1/target-variable-some-interesting-insights\">https://www.kaggle.com/alijs1/target-variable-some-interesting-insights</a></p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 344560,
      "author_name": "Wesam Elshamy",
      "author_url": "",
      "post_date": "2018-06-18T08:16:01.690000",
      "content": "<p>@alijs @AhmetErdem There is more evidence that the deal probability was generated at least partially by a rule based system.  Take a look at the deal probability distribution for each <code>param_2</code> category (<a href=\"https://www.kaggle.com/wesamelshamy/the-good-deal-feature\">plots at the bottom of this kernel</a>), you will find the following:</p>\n\n<ul>\n<li>There is a big gap between 0 and ~0.1 values</li>\n<li>About 80% of the ads have exactly 0 deal probability value</li>\n<li>As you both mentioned, the deal probability values are discrete</li>\n</ul>\n\n<p>How about training a zero deal probability model?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 343799,
      "author_name": "A_b",
      "author_url": "",
      "post_date": "2018-06-16T06:45:26.760000",
      "content": "<p>len(train['deal_probability'].unique())=18407  </p>",
      "votes": 1,
      "replies": [
        {
          "id": 344563,
          "author_name": "Lem Ko",
          "author_url": "",
          "post_date": "2018-06-18T08:19:37.963000",
          "content": "<p>More succinctly: <code>train['deal_probability'].nunique()</code></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 345870,
      "author_name": "Andrea Rapuzzi",
      "author_url": "",
      "post_date": "2018-06-20T16:21:56.257000",
      "content": "<p>It seems to be a reverse engineering problem rather than a prediction one :-)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 345872,
          "author_name": "Wesam Elshamy",
          "author_url": "",
          "post_date": "2018-06-20T16:25:55.577000",
          "content": "<p>Mostly, yes.  We are building a model that can predict what Avito model predicts.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 345875,
          "author_name": "Andrea Rapuzzi",
          "author_url": "",
          "post_date": "2018-06-20T16:29:18.407000",
          "content": "<p>So, if you have a very brilliant insight, but they had not in their model, it will render useless...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 345900,
          "author_name": "Steinhafen",
          "author_url": "",
          "post_date": "2018-06-20T17:26:00.533000",
          "content": "<p>It seems to be that the features are rated before joining the pipeline for classification.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 344398,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-06-17T21:45:58.447000",
      "content": "<p>Given they are pretty discrete, would it make sense to train some kind of classification models and use output as features?</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "343651": "![enter image description here][1]\n\nDeal probability values are very discrete. Therefore I don't think they are simple likelihoods. There must be some parent_category related parameter in its calculation because most of the values only exist for single parent category. Also the fact that nobody explains what exactly deal probability is makes the situation even more suspicious. And it seems a global minmax scaling is applied on them.\n\n\n  [1]: https://i.imgur.com/KpmdQeU.png",
    "344543": "There are some more \"strange findings\" hidden inside deal probability variable and waiting to be revealed ;) \n\nTried to show part of them in a small notebook here: https://www.kaggle.com/alijs1/target-variable-some-interesting-insights",
    "344560": "@alijs @AhmetErdem There is more evidence that the deal probability was generated at least partially by a rule based system.  Take a look at the deal probability distribution for each `param_2` category ([plots at the bottom of this kernel][1]), you will find the following:\n\n- There is a big gap between 0 and ~0.1 values\n- About 80% of the ads have exactly 0 deal probability value\n- As you both mentioned, the deal probability values are discrete\n\nHow about training a zero deal probability model?\n\n[1]: https://www.kaggle.com/wesamelshamy/the-good-deal-feature",
    "343799": "len(train['deal_probability'].unique())=18407  ",
    "345870": "It seems to be a reverse engineering problem rather than a prediction one :-)",
    "344398": "Given they are pretty discrete, would it make sense to train some kind of classification models and use output as features?"
  }
}