{
  "id": 562808,
  "title": "Kindly provide a code snippet for the custom eval-metric",
  "url": "/competitions/nexar-collision-prediction/discussion/562808",
  "author_name": "Ravi Ramakrishnan",
  "post_date": "2025-02-13T13:20:23.447000",
  "votes": 8,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hello <a href=\"https://www.kaggle.com/danielcmoura\" target=\"_blank\">@danielcmoura</a>, <a href=\"https://www.kaggle.com/orlyzvitia\" target=\"_blank\">@orlyzvitia</a>, <a href=\"https://www.kaggle.com/shizhanzhu\" target=\"_blank\">@shizhanzhu</a></p>\n<p>The evaluation metric here is a custom metric with 3 PR curves and mean precision score over the event in 500ms, 1000ms and 1500ms for accident.<br>\nCan you kindly wind this in a code snippet and provide it as it will orient everyone accordingly and create a uniform model evaluation?</p>",
  "messages": [
    {
      "id": 3123187,
      "postDate": "2025-02-13T13:20:23.447Z",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/danielcmoura\" target=\"_blank\">@danielcmoura</a>, <a href=\"https://www.kaggle.com/orlyzvitia\" target=\"_blank\">@orlyzvitia</a>, <a href=\"https://www.kaggle.com/shizhanzhu\" target=\"_blank\">@shizhanzhu</a></p>\n<p>The evaluation metric here is a custom metric with 3 PR curves and mean precision score over the event in 500ms, 1000ms and 1500ms for accident.<br>\nCan you kindly wind this in a code snippet and provide it as it will orient everyone accordingly and create a uniform model evaluation?</p>",
      "rawMarkdown": "Hello @danielcmoura, @orlyzvitia, @shizhanzhu\n\nThe evaluation metric here is a custom metric with 3 PR curves and mean precision score over the event in 500ms, 1000ms and 1500ms for accident.\nCan you kindly wind this in a code snippet and provide it as it will orient everyone accordingly and create a uniform model evaluation?",
      "votes": 8
    },
    {
      "id": 3123341,
      "postDate": "2025-02-13T16:56:07.960Z",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a>.</p>\n<p>In the end, you may find the code of the metric defined for this kaggle competition.<br>\nOn the test set, each video has a Time To Accident (TTA), meaning a time before the collision/near-miss happens (in positive cases) or a random time on normal driving. Unlike in the train set, in the test set, the video ends TTA seconds before the event. You should provide a confidence score that tells how likely a collision/near-miss is about to happen. The TTAs we use are 500ms, 1000ms and 1500ms. We only include videos that are cut at 1000ms and 1500ms TTAs if the difference between the event time and alert time is higher than 1000ms and 1500ms, respectively. </p>\n<p>If you want to use the exact same metric during your training, you would need to build a dataframe where you add the predictions of your model at different TTAs so that you can apply the metric.  </p>\n<pre><code>import numpy as \nimport pandas as pd\nimport pandas.api.types\n\nimport sklearn.metrics\n\nclass ParticipantVisibleError(Exception):\n    pass\n\ndef score(solution: pd.DataFrame, submission: pd.DataFrame, row_id_column_name: str, group_column_name: str = ) -&gt; :\n    '''\n    Mean of the Average Precision AP. AP  calculated  each grouped by wrapping\n    https://scikit-learn.org/stable/modules/generated/sklearn.metrics.average_precision_score.html\n      the  of the APs of all grouped  computed.\n\n    AP summarizes a precision-recall curve as the weighted  of precisions\n    achieved  each threshold, with the increase  recall from the previous\n    threshold used as the weight:\n\n    .. math::\n    \\text{AP} = \\sum_n (R_n - R_{n-}) P_n\n\n    where :math:`P_n`  :math:`R_n` are the precision  recall  the nth\n    threshold []. This implementation   interpolated   different\n    from computing the area under the precision-recall curve with the\n    trapezoidal rule, which uses  interpolation  can be too\n    optimistic.\n\n    Note: this implementation  restricted to the binary classification task.\n\n    Parameters\n    ----------\n    solution : ndarray of shape (n_samples,)  (n_samples, n_classes)\n    True binary   binary  indicators.\n\n    submission : ndarray of shape (n_samples,)  (n_samples, n_classes)\n    Target scores, can either be probability estimates of the positive\n    class, confidence ,  non-thresholded measure of decisions\n    (as returned by :term:`decision_function` on  classifiers).\n\n\n    Examples\n    --------\n\n    &gt;&gt;&gt; import pandas as pd\n    &gt;&gt;&gt; import numpy as \n    &gt;&gt;&gt; y_true = .([, , , ] + [,,,] + [,,,])\n    &gt;&gt;&gt; y_true = pd.DataFrame(y_true)\n    &gt;&gt;&gt; y_true[] = (len(y_true))\n    &gt;&gt;&gt; y_true[] = [, , , , , , , , , , , ]\n    &gt;&gt;&gt; y_pred = .([, , , ] * )\n    &gt;&gt;&gt; y_pred = pd.DataFrame(y_pred)\n    &gt;&gt;&gt; y_pred[] = (len(y_pred))\n    &gt;&gt;&gt; score(y_true.(), y_pred.(), , )\n    \n    '''\n\n    # Skip sorting  equality checks  the row_id_column since that should already be handled\n     solution[row_id_column_name]\n     submission[row_id_column_name]\n\n      group_column_name  solution.:\n        raise ParticipantVisibleError('Missing group column  solution')\n\n    group = solution[group_column_name]\n     solution[group_column_name]\n    groups = group.()\n\n     ((len(submission.) == )  (len(submission.) == len(solution.))):\n        raise ParticipantVisibleError(f'Invalid number of submission . Found {len(submission.)}')\n\n      pandas.api.types.is_numeric_dtype(submission.):\n        bad_dtypes = {x: submission[x].dtype   x  submission.   pandas.api.types.is_numeric_dtype(submission[x])}\n        raise ParticipantVisibleError(f'Invalid submission data types found: {bad_dtypes}')\n\n     submission.().() &gt;   submission.().() &lt; :\n        raise ParticipantVisibleError('Submitted  were  valid probabilities')\n\n    solution = solution.\n    submission = submission.\n\n    score_result = .([\n        sklearn.metrics.average_precision_score(solution[group == g], submission[group == g])\n         g  groups\n    ])\n\n     score_result\n</code></pre>",
      "rawMarkdown": "Hello @ravi20076.\n\nIn the end, you may find the code of the metric defined for this kaggle competition.\nOn the test set, each video has a Time To Accident (TTA), meaning a time before the collision/near-miss happens (in positive cases) or a random time on normal driving. Unlike in the train set, in the test set, the video ends TTA seconds before the event. You should provide a confidence score that tells how likely a collision/near-miss is about to happen. The TTAs we use are 500ms, 1000ms and 1500ms. We only include videos that are cut at 1000ms and 1500ms TTAs if the difference between the event time and alert time is higher than 1000ms and 1500ms, respectively. \n\nIf you want to use the exact same metric during your training, you would need to build a dataframe where you add the predictions of your model at different TTAs so that you can apply the metric.  \n\n```\nimport numpy as np\nimport pandas as pd\nimport pandas.api.types\n\nimport sklearn.metrics\n\nclass ParticipantVisibleError(Exception):\n    pass\n\ndef score(solution: pd.DataFrame, submission: pd.DataFrame, row_id_column_name: str, group_column_name: str = \"group\") -> float:\n    '''\n    Mean of the Average Precision AP. AP is calculated for each grouped by wrapping\n    https://scikit-learn.org/stable/modules/generated/sklearn.metrics.average_precision_score.html\n    and then the mean of the APs of all grouped is computed.\n\n    AP summarizes a precision-recall curve as the weighted mean of precisions\n    achieved at each threshold, with the increase in recall from the previous\n    threshold used as the weight:\n\n    .. math::\n    \\text{AP} = \\sum_n (R_n - R_{n-1}) P_n\n\n    where :math:`P_n` and :math:`R_n` are the precision and recall at the nth\n    threshold [1]_. This implementation is not interpolated and is different\n    from computing the area under the precision-recall curve with the\n    trapezoidal rule, which uses linear interpolation and can be too\n    optimistic.\n\n    Note: this implementation is restricted to the binary classification task.\n\n    Parameters\n    ----------\n    solution : ndarray of shape (n_samples,) or (n_samples, n_classes)\n    True binary labels or binary label indicators.\n\n    submission : ndarray of shape (n_samples,) or (n_samples, n_classes)\n    Target scores, can either be probability estimates of the positive\n    class, confidence values, or non-thresholded measure of decisions\n    (as returned by :term:`decision_function` on some classifiers).\n\n   \n    Examples\n    --------\n\n    >>> import pandas as pd\n    >>> import numpy as np\n    >>> y_true = np.array([1, 0, 0, 0] + [1,0,0,1] + [1,0,1,1])\n    >>> y_true = pd.DataFrame(y_true)\n    >>> y_true[\"id\"] = range(len(y_true))\n    >>> y_true[\"group\"] = [\"a\", \"a\", \"a\", \"a\", \"b\", \"b\", \"b\", \"b\", \"c\", \"c\", \"c\", \"c\"]\n    >>> y_pred = np.array([0.1, 0.4, 0.35, 0.8] * 3)\n    >>> y_pred = pd.DataFrame(y_pred)\n    >>> y_pred[\"id\"] = range(len(y_pred))\n    >>> score(y_true.copy(), y_pred.copy(), \"id\", \"group\")\n    0.6018518518518519\n    '''\n    \n    # Skip sorting and equality checks for the row_id_column since that should already be handled\n    del solution[row_id_column_name]\n    del submission[row_id_column_name]\n\n    if not group_column_name in solution.columns:\n        raise ParticipantVisibleError('Missing group column in solution')\n    \n    group = solution[group_column_name]\n    del solution[group_column_name]\n    groups = group.unique()\n        \n    if not((len(submission.columns) == 1) or (len(submission.columns) == len(solution.columns))):\n        raise ParticipantVisibleError(f'Invalid number of submission columns. Found {len(submission.columns)}')\n\n    if not pandas.api.types.is_numeric_dtype(submission.values):\n        bad_dtypes = {x: submission[x].dtype  for x in submission.columns if not pandas.api.types.is_numeric_dtype(submission[x])}\n        raise ParticipantVisibleError(f'Invalid submission data types found: {bad_dtypes}')\n\n    if submission.max().max() > 1 or submission.min().min() < 0:\n        raise ParticipantVisibleError('Submitted values were not valid probabilities')\n\n    solution = solution.values\n    submission = submission.values\n\n    score_result = np.mean([\n        sklearn.metrics.average_precision_score(solution[group == g], submission[group == g])\n        for g in groups\n    ])\n\n    return score_result\n```",
      "votes": 5,
      "replies": [
        {
          "id": 3123855,
          "postDate": "2025-02-14T09:29:49.870Z",
          "content": "<p>Thanks for providing such an interesting competition and also the explanation here!</p>\n<blockquote>\n  <p>We only include videos that are cut at 1000ms and 1500ms TTAs if the difference between the event time and alert time is higher than 1000ms and 1500ms, respectively.</p>\n</blockquote>\n<p>This is a very useful info, thanks for sharing!</p>",
          "rawMarkdown": "Thanks for providing such an interesting competition and also the explanation here!\n\n>We only include videos that are cut at 1000ms and 1500ms TTAs if the difference between the event time and alert time is higher than 1000ms and 1500ms, respectively.\n\nThis is a very useful info, thanks for sharing!",
          "votes": 2
        },
        {
          "id": 3133696,
          "postDate": "2025-02-25T16:08:15.757Z",
          "content": "<p>Hello. Thanks for sharing. But I have doubts regarding the implementation to training. If I understand correctly, each group has its own negatives. We can assign a group to the positives since we can cut at the right times, but how can we assign group to negatives? Does that mean there are negative groups in test (obviously unknown) and we should arbitrarily assign groups to our train negatives?</p>",
          "rawMarkdown": "Hello. Thanks for sharing. But I have doubts regarding the implementation to training. If I understand correctly, each group has its own negatives. We can assign a group to the positives since we can cut at the right times, but how can we assign group to negatives? Does that mean there are negative groups in test (obviously unknown) and we should arbitrarily assign groups to our train negatives?",
          "votes": 1,
          "replies": [
            {
              "id": 3133839,
              "postDate": "2025-02-25T17:51:24.193Z",
              "content": "<p>Hey. Thank you for your question! </p>\n<p>One note is that you can implement different metrics on your training. It makes sense to use the same metric we use on the evaluation, but basically we are tying to maximize both the Average Precision and the delta time between prediction and accident. So, if you wish, you can start with something simple that serves as a proxy for the evaluation metric. </p>\n<p>We are writing a paper that we will publish soon that explains the dataset and testing method. I attached an excerpt of the current draft of the paper that explains how we generate the test set, which I believe addresses your question. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1530400%2F9bcb1df15530710c0ad9808f52ac1094%2FScreenshot%202025-02-25%20at%2017.44.43.png?generation=1740505756175918&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "Hey. Thank you for your question! \n\nOne note is that you can implement different metrics on your training. It makes sense to use the same metric we use on the evaluation, but basically we are tying to maximize both the Average Precision and the delta time between prediction and accident. So, if you wish, you can start with something simple that serves as a proxy for the evaluation metric. \n \nWe are writing a paper that we will publish soon that explains the dataset and testing method. I attached an excerpt of the current draft of the paper that explains how we generate the test set, which I believe addresses your question. \n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1530400%2F9bcb1df15530710c0ad9808f52ac1094%2FScreenshot%202025-02-25%20at%2017.44.43.png?generation=1740505756175918&alt=media)",
              "votes": 2
            },
            {
              "id": 3133853,
              "postDate": "2025-02-25T18:00:41.017Z",
              "content": "<p>\"For negative examples, fake event times were generated by adding Gaussian noise to half of the duration of the video\" Thats very intereasting and informative for training strategies.</p>",
              "rawMarkdown": "\"For negative examples, fake event times were generated by adding Gaussian noise to half of the duration of the video\" Thats very intereasting and informative for training strategies.",
              "votes": 2
            },
            {
              "id": 3134424,
              "postDate": "2025-02-26T10:21:23.833Z",
              "content": "<p>You write \"A video is only cropped at a given Delta_tta &gt; 500 when the difference between the annotated event and alert is <strong>lower</strong> than the correspondent Delta_tta\" … I think you meant <strong>higher</strong>, since cropping a video at Delta_tta = 1000 when the difference between time_of_alert and time_of_event is lower than 1000 wouldn't allow the model to learn. Or am I missing something?</p>",
              "rawMarkdown": "You write \"A video is only cropped at a given Delta_tta > 500 when the difference between the annotated event and alert is **lower** than the correspondent Delta_tta\" ... I think you meant **higher**, since cropping a video at Delta_tta = 1000 when the difference between time_of_alert and time_of_event is lower than 1000 wouldn't allow the model to learn. Or am I missing something?",
              "votes": 1,
              "isDeleted": true
            },
            {
              "id": 3141163,
              "postDate": "2025-03-05T09:31:28.420Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 3141164,
              "postDate": "2025-03-05T09:32:01.713Z",
              "content": "<p>You are right, that was a typo, thank you!</p>",
              "rawMarkdown": "You are right, that was a typo, thank you!"
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3123341,
      "author_name": "Daniel Moura",
      "author_url": "",
      "post_date": "2025-02-13T16:56:07.960000",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a>.</p>\n<p>In the end, you may find the code of the metric defined for this kaggle competition.<br>\nOn the test set, each video has a Time To Accident (TTA), meaning a time before the collision/near-miss happens (in positive cases) or a random time on normal driving. Unlike in the train set, in the test set, the video ends TTA seconds before the event. You should provide a confidence score that tells how likely a collision/near-miss is about to happen. The TTAs we use are 500ms, 1000ms and 1500ms. We only include videos that are cut at 1000ms and 1500ms TTAs if the difference between the event time and alert time is higher than 1000ms and 1500ms, respectively. </p>\n<p>If you want to use the exact same metric during your training, you would need to build a dataframe where you add the predictions of your model at different TTAs so that you can apply the metric.  </p>\n<pre><code>import numpy as \nimport pandas as pd\nimport pandas.api.types\n\nimport sklearn.metrics\n\nclass ParticipantVisibleError(Exception):\n    pass\n\ndef score(solution: pd.DataFrame, submission: pd.DataFrame, row_id_column_name: str, group_column_name: str = ) -&gt; :\n    '''\n    Mean of the Average Precision AP. AP  calculated  each grouped by wrapping\n    https://scikit-learn.org/stable/modules/generated/sklearn.metrics.average_precision_score.html\n      the  of the APs of all grouped  computed.\n\n    AP summarizes a precision-recall curve as the weighted  of precisions\n    achieved  each threshold, with the increase  recall from the previous\n    threshold used as the weight:\n\n    .. math::\n    \\text{AP} = \\sum_n (R_n - R_{n-}) P_n\n\n    where :math:`P_n`  :math:`R_n` are the precision  recall  the nth\n    threshold []. This implementation   interpolated   different\n    from computing the area under the precision-recall curve with the\n    trapezoidal rule, which uses  interpolation  can be too\n    optimistic.\n\n    Note: this implementation  restricted to the binary classification task.\n\n    Parameters\n    ----------\n    solution : ndarray of shape (n_samples,)  (n_samples, n_classes)\n    True binary   binary  indicators.\n\n    submission : ndarray of shape (n_samples,)  (n_samples, n_classes)\n    Target scores, can either be probability estimates of the positive\n    class, confidence ,  non-thresholded measure of decisions\n    (as returned by :term:`decision_function` on  classifiers).\n\n\n    Examples\n    --------\n\n    &gt;&gt;&gt; import pandas as pd\n    &gt;&gt;&gt; import numpy as \n    &gt;&gt;&gt; y_true = .([, , , ] + [,,,] + [,,,])\n    &gt;&gt;&gt; y_true = pd.DataFrame(y_true)\n    &gt;&gt;&gt; y_true[] = (len(y_true))\n    &gt;&gt;&gt; y_true[] = [, , , , , , , , , , , ]\n    &gt;&gt;&gt; y_pred = .([, , , ] * )\n    &gt;&gt;&gt; y_pred = pd.DataFrame(y_pred)\n    &gt;&gt;&gt; y_pred[] = (len(y_pred))\n    &gt;&gt;&gt; score(y_true.(), y_pred.(), , )\n    \n    '''\n\n    # Skip sorting  equality checks  the row_id_column since that should already be handled\n     solution[row_id_column_name]\n     submission[row_id_column_name]\n\n      group_column_name  solution.:\n        raise ParticipantVisibleError('Missing group column  solution')\n\n    group = solution[group_column_name]\n     solution[group_column_name]\n    groups = group.()\n\n     ((len(submission.) == )  (len(submission.) == len(solution.))):\n        raise ParticipantVisibleError(f'Invalid number of submission . Found {len(submission.)}')\n\n      pandas.api.types.is_numeric_dtype(submission.):\n        bad_dtypes = {x: submission[x].dtype   x  submission.   pandas.api.types.is_numeric_dtype(submission[x])}\n        raise ParticipantVisibleError(f'Invalid submission data types found: {bad_dtypes}')\n\n     submission.().() &gt;   submission.().() &lt; :\n        raise ParticipantVisibleError('Submitted  were  valid probabilities')\n\n    solution = solution.\n    submission = submission.\n\n    score_result = .([\n        sklearn.metrics.average_precision_score(solution[group == g], submission[group == g])\n         g  groups\n    ])\n\n     score_result\n</code></pre>",
      "votes": 5,
      "replies": [
        {
          "id": 3123855,
          "author_name": "SLi",
          "author_url": "",
          "post_date": "2025-02-14T09:29:49.870000",
          "content": "<p>Thanks for providing such an interesting competition and also the explanation here!</p>\n<blockquote>\n  <p>We only include videos that are cut at 1000ms and 1500ms TTAs if the difference between the event time and alert time is higher than 1000ms and 1500ms, respectively.</p>\n</blockquote>\n<p>This is a very useful info, thanks for sharing!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 3133696,
          "author_name": "Ángel Jacinto Sánchez Ruiz",
          "author_url": "",
          "post_date": "2025-02-25T16:08:15.757000",
          "content": "<p>Hello. Thanks for sharing. But I have doubts regarding the implementation to training. If I understand correctly, each group has its own negatives. We can assign a group to the positives since we can cut at the right times, but how can we assign group to negatives? Does that mean there are negative groups in test (obviously unknown) and we should arbitrarily assign groups to our train negatives?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3133839,
              "author_name": "Daniel Moura",
              "author_url": "",
              "post_date": "2025-02-25T17:51:24.193000",
              "content": "<p>Hey. Thank you for your question! </p>\n<p>One note is that you can implement different metrics on your training. It makes sense to use the same metric we use on the evaluation, but basically we are tying to maximize both the Average Precision and the delta time between prediction and accident. So, if you wish, you can start with something simple that serves as a proxy for the evaluation metric. </p>\n<p>We are writing a paper that we will publish soon that explains the dataset and testing method. I attached an excerpt of the current draft of the paper that explains how we generate the test set, which I believe addresses your question. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1530400%2F9bcb1df15530710c0ad9808f52ac1094%2FScreenshot%202025-02-25%20at%2017.44.43.png?generation=1740505756175918&amp;alt=media\" alt=\"\"></p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3133853,
              "author_name": "Ángel Jacinto Sánchez Ruiz",
              "author_url": "",
              "post_date": "2025-02-25T18:00:41.017000",
              "content": "<p>\"For negative examples, fake event times were generated by adding Gaussian noise to half of the duration of the video\" Thats very intereasting and informative for training strategies.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3134424,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-02-26T10:21:23.833000",
              "content": "<p>You write \"A video is only cropped at a given Delta_tta &gt; 500 when the difference between the annotated event and alert is <strong>lower</strong> than the correspondent Delta_tta\" … I think you meant <strong>higher</strong>, since cropping a video at Delta_tta = 1000 when the difference between time_of_alert and time_of_event is lower than 1000 wouldn't allow the model to learn. Or am I missing something?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3141163,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-03-05T09:31:28.420000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3141164,
              "author_name": "Daniel Moura",
              "author_url": "",
              "post_date": "2025-03-05T09:32:01.713000",
              "content": "<p>You are right, that was a typo, thank you!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3123187": "Hello @danielcmoura, @orlyzvitia, @shizhanzhu\n\nThe evaluation metric here is a custom metric with 3 PR curves and mean precision score over the event in 500ms, 1000ms and 1500ms for accident.\nCan you kindly wind this in a code snippet and provide it as it will orient everyone accordingly and create a uniform model evaluation?",
    "3123341": "Hello @ravi20076.\n\nIn the end, you may find the code of the metric defined for this kaggle competition.\nOn the test set, each video has a Time To Accident (TTA), meaning a time before the collision/near-miss happens (in positive cases) or a random time on normal driving. Unlike in the train set, in the test set, the video ends TTA seconds before the event. You should provide a confidence score that tells how likely a collision/near-miss is about to happen. The TTAs we use are 500ms, 1000ms and 1500ms. We only include videos that are cut at 1000ms and 1500ms TTAs if the difference between the event time and alert time is higher than 1000ms and 1500ms, respectively. \n\nIf you want to use the exact same metric during your training, you would need to build a dataframe where you add the predictions of your model at different TTAs so that you can apply the metric.  \n\n```\nimport numpy as np\nimport pandas as pd\nimport pandas.api.types\n\nimport sklearn.metrics\n\nclass ParticipantVisibleError(Exception):\n    pass\n\ndef score(solution: pd.DataFrame, submission: pd.DataFrame, row_id_column_name: str, group_column_name: str = \"group\") -> float:\n    '''\n    Mean of the Average Precision AP. AP is calculated for each grouped by wrapping\n    https://scikit-learn.org/stable/modules/generated/sklearn.metrics.average_precision_score.html\n    and then the mean of the APs of all grouped is computed.\n\n    AP summarizes a precision-recall curve as the weighted mean of precisions\n    achieved at each threshold, with the increase in recall from the previous\n    threshold used as the weight:\n\n    .. math::\n    \\text{AP} = \\sum_n (R_n - R_{n-1}) P_n\n\n    where :math:`P_n` and :math:`R_n` are the precision and recall at the nth\n    threshold [1]_. This implementation is not interpolated and is different\n    from computing the area under the precision-recall curve with the\n    trapezoidal rule, which uses linear interpolation and can be too\n    optimistic.\n\n    Note: this implementation is restricted to the binary classification task.\n\n    Parameters\n    ----------\n    solution : ndarray of shape (n_samples,) or (n_samples, n_classes)\n    True binary labels or binary label indicators.\n\n    submission : ndarray of shape (n_samples,) or (n_samples, n_classes)\n    Target scores, can either be probability estimates of the positive\n    class, confidence values, or non-thresholded measure of decisions\n    (as returned by :term:`decision_function` on some classifiers).\n\n   \n    Examples\n    --------\n\n    >>> import pandas as pd\n    >>> import numpy as np\n    >>> y_true = np.array([1, 0, 0, 0] + [1,0,0,1] + [1,0,1,1])\n    >>> y_true = pd.DataFrame(y_true)\n    >>> y_true[\"id\"] = range(len(y_true))\n    >>> y_true[\"group\"] = [\"a\", \"a\", \"a\", \"a\", \"b\", \"b\", \"b\", \"b\", \"c\", \"c\", \"c\", \"c\"]\n    >>> y_pred = np.array([0.1, 0.4, 0.35, 0.8] * 3)\n    >>> y_pred = pd.DataFrame(y_pred)\n    >>> y_pred[\"id\"] = range(len(y_pred))\n    >>> score(y_true.copy(), y_pred.copy(), \"id\", \"group\")\n    0.6018518518518519\n    '''\n    \n    # Skip sorting and equality checks for the row_id_column since that should already be handled\n    del solution[row_id_column_name]\n    del submission[row_id_column_name]\n\n    if not group_column_name in solution.columns:\n        raise ParticipantVisibleError('Missing group column in solution')\n    \n    group = solution[group_column_name]\n    del solution[group_column_name]\n    groups = group.unique()\n        \n    if not((len(submission.columns) == 1) or (len(submission.columns) == len(solution.columns))):\n        raise ParticipantVisibleError(f'Invalid number of submission columns. Found {len(submission.columns)}')\n\n    if not pandas.api.types.is_numeric_dtype(submission.values):\n        bad_dtypes = {x: submission[x].dtype  for x in submission.columns if not pandas.api.types.is_numeric_dtype(submission[x])}\n        raise ParticipantVisibleError(f'Invalid submission data types found: {bad_dtypes}')\n\n    if submission.max().max() > 1 or submission.min().min() < 0:\n        raise ParticipantVisibleError('Submitted values were not valid probabilities')\n\n    solution = solution.values\n    submission = submission.values\n\n    score_result = np.mean([\n        sklearn.metrics.average_precision_score(solution[group == g], submission[group == g])\n        for g in groups\n    ])\n\n    return score_result\n```"
  }
}