{
  "id": 554252,
  "title": "Notebook Exception when using `dict_to_df` function",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/554252",
  "author_name": "IAmParadox",
  "post_date": "2024-12-31T14:00:06.416000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>If anyone is using <code>dict_to_df</code> from <a href=\"https://www.kaggle.com/code/fnands/baseline-unet-train-submit/notebook\" target=\"_blank\">notebook</a>, please update it to use the following version:</p>\n<pre><code> ():\n    \n    \n    all_coords = []\n    all_labels = []\n\n    \n     label, coords  coord_dict.items():\n         (coords):\n            all_coords.append(coords)\n            all_labels.extend([label] * (coords))\n\n    \n     all_coords:\n        all_coords = np.vstack(all_coords)\n\n     (all_coords) == :\n\n        df = pd.DataFrame({\n            : experiment_name,\n            : all_labels,\n            : [],\n            : [],\n            : []\n        })\n    :\n\n        df = pd.DataFrame({\n            : experiment_name,\n            : all_labels,\n            : all_coords[:, ],\n            : all_coords[:, ],\n            : all_coords[:, ]\n        })\n\n\n     df\n</code></pre>\n<p>The original function doesn't handle for case when the number of predicted particle is 0, this leads to <code>np.vstack</code> throwing the error. </p>\n<p>I had to spent numerous hours to debug this because my pipeline was running fine both locally and while saving the notebook version, but during the submission it threw the error.</p>",
  "messages": [
    {
      "id": 3084939,
      "postDate": "2024-12-31T14:00:06.417Z",
      "content": "<p>If anyone is using <code>dict_to_df</code> from <a href=\"https://www.kaggle.com/code/fnands/baseline-unet-train-submit/notebook\" target=\"_blank\">notebook</a>, please update it to use the following version:</p>\n<pre><code> ():\n    \n    \n    all_coords = []\n    all_labels = []\n\n    \n     label, coords  coord_dict.items():\n         (coords):\n            all_coords.append(coords)\n            all_labels.extend([label] * (coords))\n\n    \n     all_coords:\n        all_coords = np.vstack(all_coords)\n\n     (all_coords) == :\n\n        df = pd.DataFrame({\n            : experiment_name,\n            : all_labels,\n            : [],\n            : [],\n            : []\n        })\n    :\n\n        df = pd.DataFrame({\n            : experiment_name,\n            : all_labels,\n            : all_coords[:, ],\n            : all_coords[:, ],\n            : all_coords[:, ]\n        })\n\n\n     df\n</code></pre>\n<p>The original function doesn't handle for case when the number of predicted particle is 0, this leads to <code>np.vstack</code> throwing the error. </p>\n<p>I had to spent numerous hours to debug this because my pipeline was running fine both locally and while saving the notebook version, but during the submission it threw the error.</p>",
      "rawMarkdown": "If anyone is using `dict_to_df` from [notebook](https://www.kaggle.com/code/fnands/baseline-unet-train-submit/notebook), please update it to use the following version:\n\n```python\ndef dict_to_df(coord_dict, experiment_name):\n    \"\"\"\n    Convert dictionary of coordinates to pandas DataFrame.\n    \n    Parameters:\n    -----------\n    coord_dict : dict\n        Dictionary where keys are labels and values are Nx3 coordinate arrays\n        \n    Returns:\n    --------\n    pd.DataFrame\n        DataFrame with columns ['x', 'y', 'z', 'label']\n    \"\"\"\n    # Create lists to store data\n    all_coords = []\n    all_labels = []\n    \n    # Process each label and its coordinates\n    for label, coords in coord_dict.items():\n        if len(coords):\n            all_coords.append(coords)\n            all_labels.extend([label] * len(coords))\n    \n    # Concatenate all coordinates\n    if all_coords:\n        all_coords = np.vstack(all_coords)\n\n    if len(all_coords) == 0:\n\n        df = pd.DataFrame({\n            'experiment': experiment_name,\n            'particle_type': all_labels,\n            'x': [],\n            'y': [],\n            'z': []\n        })\n    else:\n    \n        df = pd.DataFrame({\n            'experiment': experiment_name,\n            'particle_type': all_labels,\n            'x': all_coords[:, 0],\n            'y': all_coords[:, 1],\n            'z': all_coords[:, 2]\n        })\n\n    \n    return df\n\n```\n\nThe original function doesn't handle for case when the number of predicted particle is 0, this leads to `np.vstack` throwing the error. \n\nI had to spent numerous hours to debug this because my pipeline was running fine both locally and while saving the notebook version, but during the submission it threw the error.",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3084939": "If anyone is using `dict_to_df` from [notebook](https://www.kaggle.com/code/fnands/baseline-unet-train-submit/notebook), please update it to use the following version:\n\n```python\ndef dict_to_df(coord_dict, experiment_name):\n    \"\"\"\n    Convert dictionary of coordinates to pandas DataFrame.\n    \n    Parameters:\n    -----------\n    coord_dict : dict\n        Dictionary where keys are labels and values are Nx3 coordinate arrays\n        \n    Returns:\n    --------\n    pd.DataFrame\n        DataFrame with columns ['x', 'y', 'z', 'label']\n    \"\"\"\n    # Create lists to store data\n    all_coords = []\n    all_labels = []\n    \n    # Process each label and its coordinates\n    for label, coords in coord_dict.items():\n        if len(coords):\n            all_coords.append(coords)\n            all_labels.extend([label] * len(coords))\n    \n    # Concatenate all coordinates\n    if all_coords:\n        all_coords = np.vstack(all_coords)\n\n    if len(all_coords) == 0:\n\n        df = pd.DataFrame({\n            'experiment': experiment_name,\n            'particle_type': all_labels,\n            'x': [],\n            'y': [],\n            'z': []\n        })\n    else:\n    \n        df = pd.DataFrame({\n            'experiment': experiment_name,\n            'particle_type': all_labels,\n            'x': all_coords[:, 0],\n            'y': all_coords[:, 1],\n            'z': all_coords[:, 2]\n        })\n\n    \n    return df\n\n```\n\nThe original function doesn't handle for case when the number of predicted particle is 0, this leads to `np.vstack` throwing the error. \n\nI had to spent numerous hours to debug this because my pipeline was running fine both locally and while saving the notebook version, but during the submission it threw the error."
  }
}