AI Briefing
KO

Patent-specialized embedding model ParaEmbed 2.0 released

·2024.06.25 09:00

Key point

XLSCOUT collaborated with Hugging Face to launch ParaEmbed 2.0, optimized for patent and intellectual property (IP) analysis.

Details

XLSCOUT unveiled ParaEmbed 2.0, developed through Hugging Face's Expert Support Program. This model is a domain-specific embedding model designed to precisely analyze the specialized terminology, context, and relationships within complex patent documents.

It was fine-tuned using expert-curated, high-quality patent data, achieving a 23% improvement in accuracy compared to the previous version. This has significantly strengthened its ability to map prior art and grasp context in patents.

Key technical features are as follows:

  • Model foundation: Built by integrating and fine-tuning open-source models such as BGE-base-v1.5, Llama 2, and Mixtral 8x7B.
  • High-performance inference: Leveraging Hugging Face's Inference Endpoints and Text Embedding Inference (TEI), it achieves high throughput, processing approximately 2,700 embeddings per second.
  • Workflow optimization: Through LLM prompt engineering and fine-tuning on instruction data, it has improved the accuracy and legal consistency of patent drafting.

This solution is used to increase efficiency in IP-related tasks such as patent invalidation search, novelty review, and patent drafting.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.