{
  "id": 57572,
  "title": "Fast Python Score Function",
  "url": "/competitions/trackml-particle-identification/discussion/57572",
  "author_name": "",
  "post_date": "2018-05-25T12:59:15.201706900Z",
  "votes": 10,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Inspired by this <a href=\"https://www.kaggle.com/vicensgaitan/r-scoring-function\">R scoring function</a> from Vicens Gaitan I implemented a scoring function in Python.  On my machine it runs 3 times faster than the one from trackml package.  On Kaggle kernels it is about 4 faster than the trackml one (speedup depends on when you run the kernel...).  See this notebook for the code: <a href=\"https://www.kaggle.com/cpmpml/a-faster-python-scoring-function\">https://www.kaggle.com/cpmpml/a-faster-python-scoring-function</a></p>\n\n<p>The function is few lines of code.</p>\n\n<p>Edit: I edited the notebook to fix package import issues, it now runs fine.</p>",
  "messages": [
    {
      "id": "333587",
      "postDate": "05/25/2018 12:59:15",
      "content": "<p>Inspired by this <a href=\"https://www.kaggle.com/vicensgaitan/r-scoring-function\">R scoring function</a> from Vicens Gaitan I implemented a scoring function in Python.  On my machine it runs 3 times faster than the one from trackml package.  On Kaggle kernels it is about 4 faster than the trackml one (speedup depends on when you run the kernel...).  See this notebook for the code: <a href=\"https://www.kaggle.com/cpmpml/a-faster-python-scoring-function\">https://www.kaggle.com/cpmpml/a-faster-python-scoring-function</a></p>\n\n<p>The function is few lines of code.</p>\n\n<p>Edit: I edited the notebook to fix package import issues, it now runs fine.</p>",
      "rawMarkdown": "Inspired by this [R scoring function][1] from Vicens Gaitan I implemented a scoring function in Python.  On my machine it runs 3 times faster than the one from trackml package.  On Kaggle kernels it is about 4 faster than the trackml one (speedup depends on when you run the kernel...).  See this notebook for the code: https://www.kaggle.com/cpmpml/a-faster-python-scoring-function\n\nThe function is few lines of code.\n\nEdit: I edited the notebook to fix package import issues, it now runs fine.\n\n  [1]: https://www.kaggle.com/vicensgaitan/r-scoring-function",
      "votes": null
    },
    {
      "id": "333625",
      "postDate": "05/25/2018 14:23:17",
      "content": "<p>Python being faster than R? damn.. thats new...</p>",
      "rawMarkdown": "Python being faster than R? damn.. thats new...",
      "votes": null
    },
    {
      "id": "333628",
      "postDate": "05/25/2018 14:26:19",
      "content": "<p>trackml library is Python as well.  I don't know how my code compares to Vicens' code.</p>",
      "rawMarkdown": "trackml library is Python as well.  I don't know how my code compares to Vicens' code.",
      "votes": null
    },
    {
      "id": "333986",
      "postDate": "05/26/2018 08:04:10",
      "content": "<p>Thanks for this we'll look into this (we're at heart C++ coders really)</p>",
      "rawMarkdown": "Thanks for this we'll look into this (we're at heart C++ coders really)",
      "votes": null
    },
    {
      "id": "334068",
      "postDate": "05/26/2018 12:53:16",
      "content": "<p>Thanks.</p>\n\n<p>I come from C++ too, and getting rid of looping in favor of built in vectorized functions in Python took me a while ;)</p>\n\n<p>You code does some checks I am not performing.  Maybe looking at the end truth dataframe in my code can achieve the same, for instance look at its length.  Duplicate ids in submission would lead to a longer one.</p>",
      "rawMarkdown": "Thanks.\n\nI come from C++ too, and getting rid of looping in favor of built in vectorized functions in Python took me a while ;)\n\nYou code does some checks I am not performing.  Maybe looking at the end truth dataframe in my code can achieve the same, for instance look at its length.  Duplicate ids in submission would lead to a longer one.",
      "votes": null
    },
    {
      "id": "335489",
      "postDate": "05/29/2018 20:12:54",
      "content": "<p>I updated the notebook with a smaller code function that runs faster on my machine.</p>",
      "rawMarkdown": "I updated the notebook with a smaller code function that runs faster on my machine.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 333625,
      "author_name": "carloshuertas",
      "author_url": "",
      "post_date": "05/25/2018 14:23:17",
      "content": "<p>Python being faster than R? damn.. thats new...</p>",
      "votes": null,
      "replies": [
        {
          "id": 333628,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/25/2018 14:26:19",
          "content": "<p>trackml library is Python as well.  I don't know how my code compares to Vicens' code.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 333986,
      "author_name": "droussea",
      "author_url": "",
      "post_date": "05/26/2018 08:04:10",
      "content": "<p>Thanks for this we'll look into this (we're at heart C++ coders really)</p>",
      "votes": null,
      "replies": [
        {
          "id": 334068,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/26/2018 12:53:16",
          "content": "<p>Thanks.</p>\n\n<p>I come from C++ too, and getting rid of looping in favor of built in vectorized functions in Python took me a while ;)</p>\n\n<p>You code does some checks I am not performing.  Maybe looking at the end truth dataframe in my code can achieve the same, for instance look at its length.  Duplicate ids in submission would lead to a longer one.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 335489,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/29/2018 20:12:54",
          "content": "<p>I updated the notebook with a smaller code function that runs faster on my machine.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "333587": "Inspired by this [R scoring function][1] from Vicens Gaitan I implemented a scoring function in Python.  On my machine it runs 3 times faster than the one from trackml package.  On Kaggle kernels it is about 4 faster than the trackml one (speedup depends on when you run the kernel...).  See this notebook for the code: https://www.kaggle.com/cpmpml/a-faster-python-scoring-function\n\nThe function is few lines of code.\n\nEdit: I edited the notebook to fix package import issues, it now runs fine.\n\n  [1]: https://www.kaggle.com/vicensgaitan/r-scoring-function",
    "333625": "Python being faster than R? damn.. thats new...",
    "333628": "trackml library is Python as well.  I don't know how my code compares to Vicens' code.",
    "333986": "Thanks for this we'll look into this (we're at heart C++ coders really)",
    "334068": "Thanks.\n\nI come from C++ too, and getting rid of looping in favor of built in vectorized functions in Python took me a while ;)\n\nYou code does some checks I am not performing.  Maybe looking at the end truth dataframe in my code can achieve the same, for instance look at its length.  Duplicate ids in submission would lead to a longer one.",
    "335489": "I updated the notebook with a smaller code function that runs faster on my machine."
  },
  "source": "meta"
}