{
  "id": 15228,
  "title": "Apache Spark - Logistic regression",
  "url": "/competitions/avito-context-ad-clicks/discussion/15228",
  "author_name": "",
  "post_date": "2015-07-14T09:49:01.747Z",
  "votes": 1,
  "comment_count": 8,
  "views": 3031,
  "content": "<p>Hi.\nI've tried to use Apache Spark and Scala in this competition.</p>\n\n<p>My code is here - <a href=\"https://www.kaggle.com/rootua/avito-context-ad-clicks/apache-spark-scala-logistic-regression\">https://www.kaggle.com/rootua/avito-context-ad-clicks/apache-spark-scala-logistic-regression</a></p>\n\n<p>I've used already implemented in MLlib algorithms - logistic regression with mini-batch gradient descent and L-BFGS (<a href=\"https://spark.apache.org/docs/latest/mllib-linear-methods.html\">https://spark.apache.org/docs/latest/mllib-linear-methods.html</a>).</p>\n\n<p>But the results are not very good. This example can't beat a benchmark.\nBest result that I've achieved was 0.13792 on LB (trained on all label=1 and 3% of label=0, with features hashing).</p>\n\n<p>Features hashing I've implemented in the same way, as here - <a href=\"https://www.kaggle.com/abhishek/avito-context-ad-clicks/beating-the-benchmark\">https://www.kaggle.com/abhishek/avito-context-ad-clicks/beating-the-benchmark</a></p>",
  "messages": [
    {
      "id": "84371",
      "postDate": "07/14/2015 09:49:01",
      "content": "<p>Hi.\nI've tried to use Apache Spark and Scala in this competition.</p>\n\n<p>My code is here - <a href=\"https://www.kaggle.com/rootua/avito-context-ad-clicks/apache-spark-scala-logistic-regression\">https://www.kaggle.com/rootua/avito-context-ad-clicks/apache-spark-scala-logistic-regression</a></p>\n\n<p>I've used already implemented in MLlib algorithms - logistic regression with mini-batch gradient descent and L-BFGS (<a href=\"https://spark.apache.org/docs/latest/mllib-linear-methods.html\">https://spark.apache.org/docs/latest/mllib-linear-methods.html</a>).</p>\n\n<p>But the results are not very good. This example can't beat a benchmark.\nBest result that I've achieved was 0.13792 on LB (trained on all label=1 and 3% of label=0, with features hashing).</p>\n\n<p>Features hashing I've implemented in the same way, as here - <a href=\"https://www.kaggle.com/abhishek/avito-context-ad-clicks/beating-the-benchmark\">https://www.kaggle.com/abhishek/avito-context-ad-clicks/beating-the-benchmark</a></p>",
      "rawMarkdown": "Hi.\r\nI've tried to use Apache Spark and Scala in this competition.\r\n\r\nMy code is here - https://www.kaggle.com/rootua/avito-context-ad-clicks/apache-spark-scala-logistic-regression\r\n\r\nI've used already implemented in MLlib algorithms - logistic regression with mini-batch gradient descent and L-BFGS (https://spark.apache.org/docs/latest/mllib-linear-methods.html).\r\n\r\nBut the results are not very good. This example can't beat a benchmark.\r\nBest result that I've achieved was 0.13792 on LB (trained on all label=1 and 3% of label=0, with features hashing).\r\n\r\nFeatures hashing I've implemented in the same way, as here - https://www.kaggle.com/abhishek/avito-context-ad-clicks/beating-the-benchmark",
      "votes": null
    },
    {
      "id": "84391",
      "postDate": "07/14/2015 12:57:29",
      "content": "<p>I've edited my script by removing features hashing.\n<a href=\"https://www.kaggle.com/rootua/avito-context-ad-clicks/apache-spark-scala-logistic-regression\">https://www.kaggle.com/rootua/avito-context-ad-clicks/apache-spark-scala-logistic-regression</a></p>\n\n<p>Now with training on 10% of all data it gives 0.063 on LB. \nStill can't beat the benchmark, but much better.</p>",
      "rawMarkdown": "I've edited my script by removing features hashing.\r\nhttps://www.kaggle.com/rootua/avito-context-ad-clicks/apache-spark-scala-logistic-regression\r\n\r\nNow with training on 10% of all data it gives 0.063 on LB. \r\nStill can't beat the benchmark, but much better.",
      "votes": null
    },
    {
      "id": "84396",
      "postDate": "07/14/2015 13:47:04",
      "content": "<p>Just first what I've noticed that you're using &quot;AdID&quot; as numeric variable with code &quot;part(1).toDouble&quot;, the same states for ObjectType and Position (in my opinion they're categorical). Also you'd not filtered parsedTrain with ObjectType == 3. <br>\nTry to use only HistCTR variable at first to build right pipeline and cv approach. </p>",
      "rawMarkdown": "Just first what I've noticed that you're using \"AdID\" as numeric variable with code \"part(1).toDouble\", the same states for ObjectType and Position (in my opinion they're categorical). Also you'd not filtered parsedTrain with ObjectType == 3.  \r\nTry to use only HistCTR variable at first to build right pipeline and cv approach.",
      "votes": null
    },
    {
      "id": "84405",
      "postDate": "07/14/2015 14:45:25",
      "content": "<p>A good example, I will try to use spark either. </p>",
      "rawMarkdown": "A good example, I will try to use spark either.",
      "votes": null
    },
    {
      "id": "84407",
      "postDate": "07/14/2015 14:54:52",
      "content": "<p>[quote=Oleksii Renov;84396]\nJust first what I've noticed that you're using &quot;AdID&quot; as numeric variable with code &quot;part(1).toDouble&quot;, the same states for ObjectType and Position (in my opinion they're categorical). Also you'd not filtered parsedTrain with ObjectType == 3. <br>\nTry to use only HistCTR variable at first to build right pipeline and cv approach. \n[/quote]</p>\n\n<p>You can see in my code that there is a features hashing function, which I was using before. I was trying to hash those categorical variables. </p>\n\n<p>The results was really bad - <strong>0.13792</strong> on LB.</p>\n\n<p>But then I found that there is a standard scaler in <strong>Logistic Regression with LBFGS</strong> in <strong>Apache Spark</strong>, so I've removed hashing and got <strong>0.063</strong> on LB.\nThat's strange, because scaling is not a solution for categorical variables.</p>",
      "rawMarkdown": "[quote=Oleksii Renov;84396]\r\nJust first what I've noticed that you're using \"AdID\" as numeric variable with code \"part(1).toDouble\", the same states for ObjectType and Position (in my opinion they're categorical). Also you'd not filtered parsedTrain with ObjectType == 3.  \r\nTry to use only HistCTR variable at first to build right pipeline and cv approach. \r\n[/quote]\r\n\r\nYou can see in my code that there is a features hashing function, which I was using before. I was trying to hash those categorical variables. \r\n\r\nThe results was really bad - **0.13792** on LB.\r\n\r\nBut then I found that there is a standard scaler in **Logistic Regression with LBFGS** in **Apache Spark**, so I've removed hashing and got **0.063** on LB.\r\nThat's strange, because scaling is not a solution for categorical variables.",
      "votes": null
    },
    {
      "id": "85473",
      "postDate": "07/15/2015 00:51:16",
      "content": "<p>I just glanced at the code - it appears that you were doing feature hashing incorrectly (line 71 on version 1).  Also, as you probably know, you need a better feature set to get any decent mean loss.</p>",
      "rawMarkdown": "I just glanced at the code - it appears that you were doing feature hashing incorrectly (line 71 on version 1).  Also, as you probably know, you need a better feature set to get any decent mean loss.",
      "votes": null
    },
    {
      "id": "85511",
      "postDate": "07/15/2015 09:28:06",
      "content": "<p>[quote=bekbolatov;85473]</p>\n\n<p>I just glanced at the code - it appears that you were doing feature hashing incorrectly (line 71 on version 1).  Also, as you probably know, you need a better feature set to get any decent mean loss.</p>\n\n<p>[/quote]</p>\n\n<p>What exactly is wrong in hashing algorithm?</p>\n\n<p>I understand that I need to spend more time on features engineering, but all that I want to do right now, is to reach the same accuracy as in this example - <a href=\"https://www.kaggle.com/abhishek/avito-context-ad-clicks/beating-the-benchmark\">https://www.kaggle.com/abhishek/avito-context-ad-clicks/beating-the-benchmark</a></p>",
      "rawMarkdown": "[quote=bekbolatov;85473]\r\n\r\nI just glanced at the code - it appears that you were doing feature hashing incorrectly (line 71 on version 1).  Also, as you probably know, you need a better feature set to get any decent mean loss.\r\n\r\n[/quote]\r\n\r\nWhat exactly is wrong in hashing algorithm?\r\n\r\nI understand that I need to spend more time on features engineering, but all that I want to do right now, is to reach the same accuracy as in this example - https://www.kaggle.com/abhishek/avito-context-ad-clicks/beating-the-benchmark",
      "votes": null
    },
    {
      "id": "86233",
      "postDate": "07/21/2015 04:51:25",
      "content": "<p>[quote=Vitalii Duk;84371]</p>\n\n<p>But the results are not very good. This example can't beat a benchmark.\nBest result that I've achieved was 0.13792 on LB (trained on all label=1 and 3% of label=0, with features hashing).</p>\n\n<p>[/quote]</p>\n\n<p>It's a big mistake to train your data as you did. You're not representing the distribution of the complete data. </p>",
      "rawMarkdown": "[quote=Vitalii Duk;84371]\r\n\r\nBut the results are not very good. This example can't beat a benchmark.\r\nBest result that I've achieved was 0.13792 on LB (trained on all label=1 and 3% of label=0, with features hashing).\r\n\r\n[/quote]\r\n\r\nIt's a big mistake to train your data as you did. You're not representing the distribution of the complete data.",
      "votes": null
    },
    {
      "id": "87175",
      "postDate": "07/27/2015 16:04:20",
      "content": "<p>I don't think you do it correct. I also can't find a correct way to utilize Hash with Logistic Regression in spark.</p>",
      "rawMarkdown": "I don't think you do it correct. I also can't find a correct way to utilize Hash with Logistic Regression in spark.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 84391,
      "author_name": "rootua",
      "author_url": "",
      "post_date": "07/14/2015 12:57:29",
      "content": "<p>I've edited my script by removing features hashing.\n<a href=\"https://www.kaggle.com/rootua/avito-context-ad-clicks/apache-spark-scala-logistic-regression\">https://www.kaggle.com/rootua/avito-context-ad-clicks/apache-spark-scala-logistic-regression</a></p>\n\n<p>Now with training on 10% of all data it gives 0.063 on LB. \nStill can't beat the benchmark, but much better.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 84396,
      "author_name": "oleksiirenov",
      "author_url": "",
      "post_date": "07/14/2015 13:47:04",
      "content": "<p>Just first what I've noticed that you're using &quot;AdID&quot; as numeric variable with code &quot;part(1).toDouble&quot;, the same states for ObjectType and Position (in my opinion they're categorical). Also you'd not filtered parsedTrain with ObjectType == 3. <br>\nTry to use only HistCTR variable at first to build right pipeline and cv approach. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 84405,
      "author_name": "yejiming",
      "author_url": "",
      "post_date": "07/14/2015 14:45:25",
      "content": "<p>A good example, I will try to use spark either. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 84407,
      "author_name": "rootua",
      "author_url": "",
      "post_date": "07/14/2015 14:54:52",
      "content": "<p>[quote=Oleksii Renov;84396]\nJust first what I've noticed that you're using &quot;AdID&quot; as numeric variable with code &quot;part(1).toDouble&quot;, the same states for ObjectType and Position (in my opinion they're categorical). Also you'd not filtered parsedTrain with ObjectType == 3. <br>\nTry to use only HistCTR variable at first to build right pipeline and cv approach. \n[/quote]</p>\n\n<p>You can see in my code that there is a features hashing function, which I was using before. I was trying to hash those categorical variables. </p>\n\n<p>The results was really bad - <strong>0.13792</strong> on LB.</p>\n\n<p>But then I found that there is a standard scaler in <strong>Logistic Regression with LBFGS</strong> in <strong>Apache Spark</strong>, so I've removed hashing and got <strong>0.063</strong> on LB.\nThat's strange, because scaling is not a solution for categorical variables.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 85473,
      "author_name": "bekbolatov",
      "author_url": "",
      "post_date": "07/15/2015 00:51:16",
      "content": "<p>I just glanced at the code - it appears that you were doing feature hashing incorrectly (line 71 on version 1).  Also, as you probably know, you need a better feature set to get any decent mean loss.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 85511,
      "author_name": "rootua",
      "author_url": "",
      "post_date": "07/15/2015 09:28:06",
      "content": "<p>[quote=bekbolatov;85473]</p>\n\n<p>I just glanced at the code - it appears that you were doing feature hashing incorrectly (line 71 on version 1).  Also, as you probably know, you need a better feature set to get any decent mean loss.</p>\n\n<p>[/quote]</p>\n\n<p>What exactly is wrong in hashing algorithm?</p>\n\n<p>I understand that I need to spend more time on features engineering, but all that I want to do right now, is to reach the same accuracy as in this example - <a href=\"https://www.kaggle.com/abhishek/avito-context-ad-clicks/beating-the-benchmark\">https://www.kaggle.com/abhishek/avito-context-ad-clicks/beating-the-benchmark</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 86233,
      "author_name": "quacktau",
      "author_url": "",
      "post_date": "07/21/2015 04:51:25",
      "content": "<p>[quote=Vitalii Duk;84371]</p>\n\n<p>But the results are not very good. This example can't beat a benchmark.\nBest result that I've achieved was 0.13792 on LB (trained on all label=1 and 3% of label=0, with features hashing).</p>\n\n<p>[/quote]</p>\n\n<p>It's a big mistake to train your data as you did. You're not representing the distribution of the complete data. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 87175,
      "author_name": "yejiming",
      "author_url": "",
      "post_date": "07/27/2015 16:04:20",
      "content": "<p>I don't think you do it correct. I also can't find a correct way to utilize Hash with Logistic Regression in spark.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "84371": "Hi.\r\nI've tried to use Apache Spark and Scala in this competition.\r\n\r\nMy code is here - https://www.kaggle.com/rootua/avito-context-ad-clicks/apache-spark-scala-logistic-regression\r\n\r\nI've used already implemented in MLlib algorithms - logistic regression with mini-batch gradient descent and L-BFGS (https://spark.apache.org/docs/latest/mllib-linear-methods.html).\r\n\r\nBut the results are not very good. This example can't beat a benchmark.\r\nBest result that I've achieved was 0.13792 on LB (trained on all label=1 and 3% of label=0, with features hashing).\r\n\r\nFeatures hashing I've implemented in the same way, as here - https://www.kaggle.com/abhishek/avito-context-ad-clicks/beating-the-benchmark",
    "84391": "I've edited my script by removing features hashing.\r\nhttps://www.kaggle.com/rootua/avito-context-ad-clicks/apache-spark-scala-logistic-regression\r\n\r\nNow with training on 10% of all data it gives 0.063 on LB. \r\nStill can't beat the benchmark, but much better.",
    "84396": "Just first what I've noticed that you're using \"AdID\" as numeric variable with code \"part(1).toDouble\", the same states for ObjectType and Position (in my opinion they're categorical). Also you'd not filtered parsedTrain with ObjectType == 3.  \r\nTry to use only HistCTR variable at first to build right pipeline and cv approach.",
    "84405": "A good example, I will try to use spark either.",
    "84407": "[quote=Oleksii Renov;84396]\r\nJust first what I've noticed that you're using \"AdID\" as numeric variable with code \"part(1).toDouble\", the same states for ObjectType and Position (in my opinion they're categorical). Also you'd not filtered parsedTrain with ObjectType == 3.  \r\nTry to use only HistCTR variable at first to build right pipeline and cv approach. \r\n[/quote]\r\n\r\nYou can see in my code that there is a features hashing function, which I was using before. I was trying to hash those categorical variables. \r\n\r\nThe results was really bad - **0.13792** on LB.\r\n\r\nBut then I found that there is a standard scaler in **Logistic Regression with LBFGS** in **Apache Spark**, so I've removed hashing and got **0.063** on LB.\r\nThat's strange, because scaling is not a solution for categorical variables.",
    "85473": "I just glanced at the code - it appears that you were doing feature hashing incorrectly (line 71 on version 1).  Also, as you probably know, you need a better feature set to get any decent mean loss.",
    "85511": "[quote=bekbolatov;85473]\r\n\r\nI just glanced at the code - it appears that you were doing feature hashing incorrectly (line 71 on version 1).  Also, as you probably know, you need a better feature set to get any decent mean loss.\r\n\r\n[/quote]\r\n\r\nWhat exactly is wrong in hashing algorithm?\r\n\r\nI understand that I need to spend more time on features engineering, but all that I want to do right now, is to reach the same accuracy as in this example - https://www.kaggle.com/abhishek/avito-context-ad-clicks/beating-the-benchmark",
    "86233": "[quote=Vitalii Duk;84371]\r\n\r\nBut the results are not very good. This example can't beat a benchmark.\r\nBest result that I've achieved was 0.13792 on LB (trained on all label=1 and 3% of label=0, with features hashing).\r\n\r\n[/quote]\r\n\r\nIt's a big mistake to train your data as you did. You're not representing the distribution of the complete data.",
    "87175": "I don't think you do it correct. I also can't find a correct way to utilize Hash with Logistic Regression in spark."
  },
  "source": "meta"
}