{
  "id": 15135,
  "title": "Logistic regression - scikit-learn",
  "url": "/competitions/avito-context-ad-clicks/discussion/15135",
  "author_name": "",
  "post_date": "2015-07-09T10:41:48.897Z",
  "votes": null,
  "comment_count": 1,
  "views": 955,
  "content": "<p>I'm trying to understand FTRL algorithm which was published here - <a href=\"https://www.kaggle.com/c/avito-context-ad-clicks/forums/t/14516/beating-the-benchmark\">https://www.kaggle.com/c/avito-context-ad-clicks/forums/t/14516/beating-the-benchmark</a>.</p>\n\n<p>Really I don't understand why in the example we're submitting file with probabilities, but in Evaluation (<a href=\"https://www.kaggle.com/c/avito-context-ad-clicks/details/evaluation\">https://www.kaggle.com/c/avito-context-ad-clicks/details/evaluation</a>) it says that submission file must look like:</p>\n\n<pre><code>TestId,IsClick\n0,0\n1,0\n5,0\n7,0\n9,0\netc.\n</code></pre>\n\n<p>I've tried to re-write it, and use just <strong>Logistic regression with SGD</strong> from <strong>scikit-learn</strong>.\n<strong>lr_sgd_own-hash_proba.py</strong> gives me a little bit better results than the original example, maybe because I've trained it on 16 000 000 examples. But it proves that Logistic Regression with SGD can give the same results as FTRL.</p>\n\n<p>Also when I only modify <strong>lr_sgd_own-hash_proba.py</strong> file to output <strong>1</strong> or <strong>0</strong> (just by changing <strong>predict_proba()</strong> to <strong>predict()</strong> ), not a probability it gives  <strong>0.3 LogLoss</strong> on Kaggle, after submission. Can anybody explain why it happens?</p>\n\n<p>There other thing that I don't understand, it's why hashing gives us better results than the one without hashing?\nCan anybody explain how it works and why do we need to use it here? We don't have categorical features here, all of our features are numerical (if we're using only original columns from <strong>trainSearchStream.tsv</strong> or <strong>testSearchStream.tsv</strong>).\n<strong>lr_sgd_nohash_proba.py</strong> shows extremely bad results. LogLoss is very high (1.1 on local data, 3.1 on Kaggle).</p>\n\n<p>Also I have no idea how to use <strong>FeatureHasher</strong> from <strong>scikit-learn</strong> here. It's output is little bit strange. When I pass one training example to <strong>fit_transform()</strong> method of <strong>FeatureHasher(n_features=20, input_type='string', non_negative=True)</strong> object, I get next results:</p>\n\n<pre><code>(0, 5)  0.0\n  (0, 6)    1.0\n  (0, 8)    1.0\n  (0, 10)   1.0\n  (0, 14)   1.0\n  (0, 17)   1.0\n  (0, 19)   1.0\n  (1, 2)    1.0\n  (1, 6)    1.0\n  (1, 14)   1.0\n  (1, 19)   1.0\n  (2, 1)    3.0\n  (2, 3)    2.0\n  (2, 4)    1.0\n  (2, 19)   2.0\n  (3, 1)    2.0\n  (3, 2)    1.0\n  (3, 6)    1.0\n  (3, 11)   1.0\n  (3, 14)   1.0\n  (3, 17)   1.0\n  (3, 19)   1.0\n</code></pre>\n\n<p>Don't know how can I use this in classification.</p>\n\n<p>Will appreciate for your help.</p>",
  "messages": [
    {
      "id": "83916",
      "postDate": "07/09/2015 10:41:48",
      "content": "<p>I'm trying to understand FTRL algorithm which was published here - <a href=\"https://www.kaggle.com/c/avito-context-ad-clicks/forums/t/14516/beating-the-benchmark\">https://www.kaggle.com/c/avito-context-ad-clicks/forums/t/14516/beating-the-benchmark</a>.</p>\n\n<p>Really I don't understand why in the example we're submitting file with probabilities, but in Evaluation (<a href=\"https://www.kaggle.com/c/avito-context-ad-clicks/details/evaluation\">https://www.kaggle.com/c/avito-context-ad-clicks/details/evaluation</a>) it says that submission file must look like:</p>\n\n<pre><code>TestId,IsClick\n0,0\n1,0\n5,0\n7,0\n9,0\netc.\n</code></pre>\n\n<p>I've tried to re-write it, and use just <strong>Logistic regression with SGD</strong> from <strong>scikit-learn</strong>.\n<strong>lr_sgd_own-hash_proba.py</strong> gives me a little bit better results than the original example, maybe because I've trained it on 16 000 000 examples. But it proves that Logistic Regression with SGD can give the same results as FTRL.</p>\n\n<p>Also when I only modify <strong>lr_sgd_own-hash_proba.py</strong> file to output <strong>1</strong> or <strong>0</strong> (just by changing <strong>predict_proba()</strong> to <strong>predict()</strong> ), not a probability it gives  <strong>0.3 LogLoss</strong> on Kaggle, after submission. Can anybody explain why it happens?</p>\n\n<p>There other thing that I don't understand, it's why hashing gives us better results than the one without hashing?\nCan anybody explain how it works and why do we need to use it here? We don't have categorical features here, all of our features are numerical (if we're using only original columns from <strong>trainSearchStream.tsv</strong> or <strong>testSearchStream.tsv</strong>).\n<strong>lr_sgd_nohash_proba.py</strong> shows extremely bad results. LogLoss is very high (1.1 on local data, 3.1 on Kaggle).</p>\n\n<p>Also I have no idea how to use <strong>FeatureHasher</strong> from <strong>scikit-learn</strong> here. It's output is little bit strange. When I pass one training example to <strong>fit_transform()</strong> method of <strong>FeatureHasher(n_features=20, input_type='string', non_negative=True)</strong> object, I get next results:</p>\n\n<pre><code>(0, 5)  0.0\n  (0, 6)    1.0\n  (0, 8)    1.0\n  (0, 10)   1.0\n  (0, 14)   1.0\n  (0, 17)   1.0\n  (0, 19)   1.0\n  (1, 2)    1.0\n  (1, 6)    1.0\n  (1, 14)   1.0\n  (1, 19)   1.0\n  (2, 1)    3.0\n  (2, 3)    2.0\n  (2, 4)    1.0\n  (2, 19)   2.0\n  (3, 1)    2.0\n  (3, 2)    1.0\n  (3, 6)    1.0\n  (3, 11)   1.0\n  (3, 14)   1.0\n  (3, 17)   1.0\n  (3, 19)   1.0\n</code></pre>\n\n<p>Don't know how can I use this in classification.</p>\n\n<p>Will appreciate for your help.</p>",
      "rawMarkdown": "I'm trying to understand FTRL algorithm which was published here - https://www.kaggle.com/c/avito-context-ad-clicks/forums/t/14516/beating-the-benchmark.\r\n\r\nReally I don't understand why in the example we're submitting file with probabilities, but in Evaluation (https://www.kaggle.com/c/avito-context-ad-clicks/details/evaluation) it says that submission file must look like:\r\n\r\n    TestId,IsClick\r\n    0,0\r\n    1,0\r\n    5,0\r\n    7,0\r\n    9,0\r\n    etc.\r\n\r\nI've tried to re-write it, and use just **Logistic regression with SGD** from **scikit-learn**.\r\n**lr_sgd_own-hash_proba.py** gives me a little bit better results than the original example, maybe because I've trained it on 16 000 000 examples. But it proves that Logistic Regression with SGD can give the same results as FTRL.\r\n\r\nAlso when I only modify **lr_sgd_own-hash_proba.py** file to output **1** or **0** (just by changing **predict_proba()** to **predict()** ), not a probability it gives  **0.3 LogLoss** on Kaggle, after submission. Can anybody explain why it happens?\r\n \r\nThere other thing that I don't understand, it's why hashing gives us better results than the one without hashing?\r\nCan anybody explain how it works and why do we need to use it here? We don't have categorical features here, all of our features are numerical (if we're using only original columns from **trainSearchStream.tsv** or **testSearchStream.tsv**).\r\n**lr_sgd_nohash_proba.py** shows extremely bad results. LogLoss is very high (1.1 on local data, 3.1 on Kaggle).\r\n\r\nAlso I have no idea how to use **FeatureHasher** from **scikit-learn** here. It's output is little bit strange. When I pass one training example to **fit_transform()** method of **FeatureHasher(n_features=20, input_type='string', non_negative=True)** object, I get next results:\r\n\r\n    (0, 5)\t0.0\r\n      (0, 6)\t1.0\r\n      (0, 8)\t1.0\r\n      (0, 10)\t1.0\r\n      (0, 14)\t1.0\r\n      (0, 17)\t1.0\r\n      (0, 19)\t1.0\r\n      (1, 2)\t1.0\r\n      (1, 6)\t1.0\r\n      (1, 14)\t1.0\r\n      (1, 19)\t1.0\r\n      (2, 1)\t3.0\r\n      (2, 3)\t2.0\r\n      (2, 4)\t1.0\r\n      (2, 19)\t2.0\r\n      (3, 1)\t2.0\r\n      (3, 2)\t1.0\r\n      (3, 6)\t1.0\r\n      (3, 11)\t1.0\r\n      (3, 14)\t1.0\r\n      (3, 17)\t1.0\r\n      (3, 19)\t1.0\r\n\r\nDon't know how can I use this in classification.\r\n\r\nWill appreciate for your help.",
      "votes": null
    },
    {
      "id": "83940",
      "postDate": "07/09/2015 14:29:51",
      "content": "<p>[quote=Vitalii Duk;83916]\nAlso when I only modify <strong>lr_sgd_own-hash_proba.py</strong> file to output <strong>1</strong> or <strong>0</strong> (just by changing <strong>predict_proba()</strong> to <strong>predict()</strong> ), not a probability it gives  <strong>0.3 LogLoss</strong> on Kaggle, after submission. Can anybody explain why it happens?\n[/quote]</p>\n\n<p>The Logloss metric penalize heavily for &quot;overconfidence&quot; - look at the equation again and remember that log(0) = -Infinity. So sending 0 and 1 prediction will give you very bad results for sure, unless you predict them all perfectly.</p>",
      "rawMarkdown": "[quote=Vitalii Duk;83916]\r\nAlso when I only modify **lr_sgd_own-hash_proba.py** file to output **1** or **0** (just by changing **predict_proba()** to **predict()** ), not a probability it gives  **0.3 LogLoss** on Kaggle, after submission. Can anybody explain why it happens?\r\n[/quote]\r\n\r\nThe Logloss metric penalize heavily for \"overconfidence\" - look at the equation again and remember that log(0) = -Infinity. So sending 0 and 1 prediction will give you very bad results for sure, unless you predict them all perfectly.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 83940,
      "author_name": "npetitclerc",
      "author_url": "",
      "post_date": "07/09/2015 14:29:51",
      "content": "<p>[quote=Vitalii Duk;83916]\nAlso when I only modify <strong>lr_sgd_own-hash_proba.py</strong> file to output <strong>1</strong> or <strong>0</strong> (just by changing <strong>predict_proba()</strong> to <strong>predict()</strong> ), not a probability it gives  <strong>0.3 LogLoss</strong> on Kaggle, after submission. Can anybody explain why it happens?\n[/quote]</p>\n\n<p>The Logloss metric penalize heavily for &quot;overconfidence&quot; - look at the equation again and remember that log(0) = -Infinity. So sending 0 and 1 prediction will give you very bad results for sure, unless you predict them all perfectly.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "83916": "I'm trying to understand FTRL algorithm which was published here - https://www.kaggle.com/c/avito-context-ad-clicks/forums/t/14516/beating-the-benchmark.\r\n\r\nReally I don't understand why in the example we're submitting file with probabilities, but in Evaluation (https://www.kaggle.com/c/avito-context-ad-clicks/details/evaluation) it says that submission file must look like:\r\n\r\n    TestId,IsClick\r\n    0,0\r\n    1,0\r\n    5,0\r\n    7,0\r\n    9,0\r\n    etc.\r\n\r\nI've tried to re-write it, and use just **Logistic regression with SGD** from **scikit-learn**.\r\n**lr_sgd_own-hash_proba.py** gives me a little bit better results than the original example, maybe because I've trained it on 16 000 000 examples. But it proves that Logistic Regression with SGD can give the same results as FTRL.\r\n\r\nAlso when I only modify **lr_sgd_own-hash_proba.py** file to output **1** or **0** (just by changing **predict_proba()** to **predict()** ), not a probability it gives  **0.3 LogLoss** on Kaggle, after submission. Can anybody explain why it happens?\r\n \r\nThere other thing that I don't understand, it's why hashing gives us better results than the one without hashing?\r\nCan anybody explain how it works and why do we need to use it here? We don't have categorical features here, all of our features are numerical (if we're using only original columns from **trainSearchStream.tsv** or **testSearchStream.tsv**).\r\n**lr_sgd_nohash_proba.py** shows extremely bad results. LogLoss is very high (1.1 on local data, 3.1 on Kaggle).\r\n\r\nAlso I have no idea how to use **FeatureHasher** from **scikit-learn** here. It's output is little bit strange. When I pass one training example to **fit_transform()** method of **FeatureHasher(n_features=20, input_type='string', non_negative=True)** object, I get next results:\r\n\r\n    (0, 5)\t0.0\r\n      (0, 6)\t1.0\r\n      (0, 8)\t1.0\r\n      (0, 10)\t1.0\r\n      (0, 14)\t1.0\r\n      (0, 17)\t1.0\r\n      (0, 19)\t1.0\r\n      (1, 2)\t1.0\r\n      (1, 6)\t1.0\r\n      (1, 14)\t1.0\r\n      (1, 19)\t1.0\r\n      (2, 1)\t3.0\r\n      (2, 3)\t2.0\r\n      (2, 4)\t1.0\r\n      (2, 19)\t2.0\r\n      (3, 1)\t2.0\r\n      (3, 2)\t1.0\r\n      (3, 6)\t1.0\r\n      (3, 11)\t1.0\r\n      (3, 14)\t1.0\r\n      (3, 17)\t1.0\r\n      (3, 19)\t1.0\r\n\r\nDon't know how can I use this in classification.\r\n\r\nWill appreciate for your help.",
    "83940": "[quote=Vitalii Duk;83916]\r\nAlso when I only modify **lr_sgd_own-hash_proba.py** file to output **1** or **0** (just by changing **predict_proba()** to **predict()** ), not a probability it gives  **0.3 LogLoss** on Kaggle, after submission. Can anybody explain why it happens?\r\n[/quote]\r\n\r\nThe Logloss metric penalize heavily for \"overconfidence\" - look at the equation again and remember that log(0) = -Infinity. So sending 0 and 1 prediction will give you very bad results for sure, unless you predict them all perfectly."
  },
  "source": "meta"
}