{
  "id": 515053,
  "title": "Is this real explainability?",
  "url": "/competitions/uspto-explainable-ai/discussion/515053",
  "author_name": "",
  "post_date": "2024-06-26T17:03:11.376758900Z",
  "votes": 8,
  "comment_count": 1,
  "views": 0,
  "content": "<p>First of all, I really like the problem tackled by the competition. Semantic similarity search is an amazing tool, but it can be a little opaque, especially if the model used is not known.</p>\n<p>At the same time, I can't help but wonder - what are we gaining here by building the queries? Instead of trying to understand how the similarity model thinks (explain its behaviour), we are essentially building an alternative explanation for the same outcome. As an example, think of the strange ways the ancient people tried to explain how planets move. Two alternative models, seemingly correctly explaining the same observed fenomenon, but one would never allow us to progress astronomy and space exploration.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3372786%2Fd5f1b772562d31b6274c012185ed0b07%2Fptolemy-geocentric.webp?generation=1719420410510337&amp;alt=media\" alt=\"Geocentricism with epicycles\"></p>\n<p>Wouldn't it make more sense to focus on the embedding_v1 field from the patent dataset, since it is the actual scoring model of the neighbours? Maybe try and see which embedding vector features are most similar, then try and come up with what the features mean? Maybe see if a decoder model can be trained to generate explanations based on the original patent text and/or embeddings?</p>\n<p>No disrespect meant, explainability is a hard, unsolved problem. I'm just trying to understand how this particular approach would be valuable to the patent experts, other than looking familiar.</p>\n<p>Thanks!</p>",
  "messages": [
    {
      "id": "2891513",
      "postDate": "06/26/2024 17:03:11",
      "content": "<p>First of all, I really like the problem tackled by the competition. Semantic similarity search is an amazing tool, but it can be a little opaque, especially if the model used is not known.</p>\n<p>At the same time, I can't help but wonder - what are we gaining here by building the queries? Instead of trying to understand how the similarity model thinks (explain its behaviour), we are essentially building an alternative explanation for the same outcome. As an example, think of the strange ways the ancient people tried to explain how planets move. Two alternative models, seemingly correctly explaining the same observed fenomenon, but one would never allow us to progress astronomy and space exploration.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3372786%2Fd5f1b772562d31b6274c012185ed0b07%2Fptolemy-geocentric.webp?generation=1719420410510337&amp;alt=media\" alt=\"Geocentricism with epicycles\"></p>\n<p>Wouldn't it make more sense to focus on the embedding_v1 field from the patent dataset, since it is the actual scoring model of the neighbours? Maybe try and see which embedding vector features are most similar, then try and come up with what the features mean? Maybe see if a decoder model can be trained to generate explanations based on the original patent text and/or embeddings?</p>\n<p>No disrespect meant, explainability is a hard, unsolved problem. I'm just trying to understand how this particular approach would be valuable to the patent experts, other than looking familiar.</p>\n<p>Thanks!</p>",
      "rawMarkdown": "First of all, I really like the problem tackled by the competition. Semantic similarity search is an amazing tool, but it can be a little opaque, especially if the model used is not known.\n\nAt the same time, I can't help but wonder - what are we gaining here by building the queries? Instead of trying to understand how the similarity model thinks (explain its behaviour), we are essentially building an alternative explanation for the same outcome. As an example, think of the strange ways the ancient people tried to explain how planets move. Two alternative models, seemingly correctly explaining the same observed fenomenon, but one would never allow us to progress astronomy and space exploration.\n\n![Geocentricism with epicycles](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3372786%2Fd5f1b772562d31b6274c012185ed0b07%2Fptolemy-geocentric.webp?generation=1719420410510337&alt=media)\n\nWouldn't it make more sense to focus on the embedding_v1 field from the patent dataset, since it is the actual scoring model of the neighbours? Maybe try and see which embedding vector features are most similar, then try and come up with what the features mean? Maybe see if a decoder model can be trained to generate explanations based on the original patent text and/or embeddings?\n\nNo disrespect meant, explainability is a hard, unsolved problem. I'm just trying to understand how this particular approach would be valuable to the patent experts, other than looking familiar.\n\nThanks!",
      "votes": null
    },
    {
      "id": "2926618",
      "postDate": "07/17/2024 20:46:05",
      "content": "<p>Yes, I am a bit disappointed after understanding what the competition is about. It basically asks us to TRUST that the LLM's nearest neighbor search is reliable. So, if someone builds a query that is in fact superior to the LLM's predictions, they would not win this competition.</p>",
      "rawMarkdown": "Yes, I am a bit disappointed after understanding what the competition is about. It basically asks us to TRUST that the LLM's nearest neighbor search is reliable. So, if someone builds a query that is in fact superior to the LLM's predictions, they would not win this competition.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2926618,
      "author_name": "celestiallylvd1",
      "author_url": "",
      "post_date": "07/17/2024 20:46:05",
      "content": "<p>Yes, I am a bit disappointed after understanding what the competition is about. It basically asks us to TRUST that the LLM's nearest neighbor search is reliable. So, if someone builds a query that is in fact superior to the LLM's predictions, they would not win this competition.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2891513": "First of all, I really like the problem tackled by the competition. Semantic similarity search is an amazing tool, but it can be a little opaque, especially if the model used is not known.\n\nAt the same time, I can't help but wonder - what are we gaining here by building the queries? Instead of trying to understand how the similarity model thinks (explain its behaviour), we are essentially building an alternative explanation for the same outcome. As an example, think of the strange ways the ancient people tried to explain how planets move. Two alternative models, seemingly correctly explaining the same observed fenomenon, but one would never allow us to progress astronomy and space exploration.\n\n![Geocentricism with epicycles](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3372786%2Fd5f1b772562d31b6274c012185ed0b07%2Fptolemy-geocentric.webp?generation=1719420410510337&alt=media)\n\nWouldn't it make more sense to focus on the embedding_v1 field from the patent dataset, since it is the actual scoring model of the neighbours? Maybe try and see which embedding vector features are most similar, then try and come up with what the features mean? Maybe see if a decoder model can be trained to generate explanations based on the original patent text and/or embeddings?\n\nNo disrespect meant, explainability is a hard, unsolved problem. I'm just trying to understand how this particular approach would be valuable to the patent experts, other than looking familiar.\n\nThanks!",
    "2926618": "Yes, I am a bit disappointed after understanding what the competition is about. It basically asks us to TRUST that the LLM's nearest neighbor search is reliable. So, if someone builds a query that is in fact superior to the LLM's predictions, they would not win this competition."
  },
  "source": "meta"
}